Make AI Fit
EN

Glossary term

Tokens per second

A rough way to feel whether a local model interaction will stay comfortable enough for real use.

Tokens per second is a practical performance signal that tells you how quickly a model can generate output during inference. Category: Hardware Related topics: 0 Related resources: 1

Why this term matters

For many local workflows, usability is less about peak benchmarks and more about whether the interaction stays fast enough that you do not abandon the setup after a week.

Examples

  • A box that technically runs the model may still be a poor fit if the response speed breaks the workflow.
  • Beginners often benefit more from a slightly smaller model that stays responsive than a heavier model that feels impressive but sluggish.

Pages where this term matters