The act of a model generating an answer (as opposed to training, which is how it learned). Every time you chat, the API runs one inference.