What is a small language model?
A small language model has fewer learned parameters than a larger comparison, although there is no universal size threshold. Training, architecture and the task also affect what it can do.
Size is separate from openness and hosting. A small model can be provided through a hosted service, while downloadable weights do not establish that a model will run acceptably on a company’s equipment.
A June 2025 position paper from NVIDIA and Georgia Tech argues that smaller models can handle bounded work inside agent systems. That is a research argument based on the technology considered then, rather than a current marketing benchmark. The position paper
Which marketing tasks can smaller models handle?
Smaller models can be candidates for clearly specified tasks such as extracting fields, assigning categories or reformatting approved material. Their sufficiency depends on the accuracy and consistency achieved on the actual inputs.
These tasks have relatively explicit outputs to inspect. Interpreting contradictory customer evidence or developing a strategic position involves broader judgement and may benefit from a more capable model. Size alone cannot determine that outcome.
A workflow can also use different models for different tasks. Routing simple work to one model and uncertain work elsewhere introduces its own maintenance and evaluation requirements. A mixed system is an operating choice, rather than an automatic saving.
Can small models replace 80% of large-model work?
The cited research does not establish that 80% of marketing work can move to small models. Its appendix estimates replaceable query shares in three selected agent designs, not a representative marketing function. The paper and appendix
A query is also different from a completed job. One assignment may contain several easy steps and a consequential decision that remains difficult. Counting the easy calls does not establish that the entire assignment is equally replaceable.
The paper is more than a year old. Changes in model capability, deployment and pricing mean its comparisons cannot rank the systems available today or set a current budget assumption.
When does a smaller model reduce the full cost?
A smaller model reduces the full cost when any processing saving exceeds additional correction, review and operating costs at the required quality. Lower usage prices alone do not establish that saving.
Cost per accepted output includes failed attempts and work redirected to another model or person. The cost of maintaining a routing system also affects the comparison, particularly when volumes are low or tasks change frequently.
A relevant comparison uses the same input material and acceptance criteria for each option. Results remain specific to that task and model version. Measuring return on AI covers the distinction between cheaper generation and a better overall result.