Inference time compute: what happens when an AI model answers
Inference is the work a trained AI model performs when it receives a new input and produces an output. It is distinct from training, and it is where the size of a request, the desired quality and the available capacity begin to shape the experience of using an AI system.





