Small models might be an answer, with work, but should they be?

I pull research once a week on nine topics, one of which is small models and LoRA to see what agrees, disagrees or provides new ideas of exploration. I have recently been digging deeper into how smaller models understand sentences and how to trigger deterministic code from them making sense of meaning. In my curated pull of research (AI Assisted), I found a paper comparing 1B to 2B models against 7B to 8B models at detecting GPS spoofing attacks on vehicles. Code first turned the GPS and other sensor readings into written descriptions, and a model then sorted each description into one of five labels. The purpose of the paper was comparison of model capability, not if using models was the best fit for the task. There was no comparison against code checking the sensor readings directly.
There are numerous choices for models to do work now at different types of capability and size. I continue to ask the question where do small models fit, but my answer to this continues to narrow. This happens because when you need to propose a solution as a builder or architect the real question is what is the best component to do the work. That component can be a model or deterministic code. The overhead of evals, security and monitoring of behavioral systems tends to tip the scale in favor of deterministic code in the solutions I am considering. Small models can still be considered when cost is a factor or the data has to stay on premises.
During testing some categories have come up. They come from quantized models doing one kind of job, pulling structured facts out of a written client description against a strict JSON schema.
- Under 2B parameters nothing worked. With a defined JSON schema the model never stopped writing and ran out its token limit on all 170 tries. Zero valid.
- Between 2B and 9B the models kept the JSON format on 65% to 85% of tries. They got the task right on 2 to 10 of 34 cases.
- The 14B model, from an older generation, did worse than the 9B. It kept the format on about half its tries, got 8 of 34 right, and gave a different answer to the same input on almost half the cases.
- The 27B to 30B models kept the format on 165 to 170 of 170 tries. The 30B was a mixture of experts, which uses about 3B of its parameters for each word it writes. They got the task right on 10 to 12 of 34 cases. Models this size adhered to the JSON format, but usually did not produce the correct answer.
Chaining models together is a separate problem from model size. Each component can work alone and still fail once they are integrated, as problems propagate down the line. I had a discussion about enhancing customer service client satisfaction. The desire was to have AI analyze the live conversation, determine the next best possible answer and do sentiment analysis. Each one of those items separately can make sense for AI to perform, but chain the system together with the understanding of what models can do reliably and you have a problem.
- Speech to text is a good use of a model.
- Classification within determined bounds is a model capability, when the input is open speech or text that code cannot sort on its own.
- Retrieval of data whether through RAG or tool calls is done regularly with models.
- Well described tool calls with clear failure flows work with models.
- Structured output is doable with a defined format and is a best practice with tool calls.
- Post call analysis of a transcript is a capability I have built before for customer service, including sentiment analysis.
Most of items 1 through 5 have to run sequentially while the customer is still talking. Item 6 has a better chance because it runs after the call ends and can take longer to finish. The representative would most likely have delayed alerts and this still assumes classification, tool calls and adherence to a schema is accurate. To attempt this you should have deterministic code surrounding a model where they can do the most good. I would start with post call analysis and work backwards to solve each of the capabilities until I can process a live call within time constraints.
All of this becomes more problematic if you were to require audit capture and designate the conversation to have protected data status like in HIPAA (healthcare related). Experience has also taught me there is a difference between just capturing Chain Of Thought (COT) to show what the model was thinking vs. an audit that grounds to data of why the model thought what it did. Grounding of answers is far more difficult than capturing COT. My "aha" moment came to me when I realized a hallucinated answer can be explained with hallucinated thinking or COT.
Trying to apply LLM capabilities to solutions has also taught me their limitations. My rule of thumb is currently use deterministic code wherever possible, and use a model only where code cannot do the job. Where I do plan to use one, a person or code checks what it returns. The reality is even when they are used the models I tested, up to 30B, will require considerable support from surrounding code to be effective.
Research
Across four models, 69% of annotated chain-of-thought unfaithfulness occurred on incorrect answers, and on those incorrect answers no tested detection signal performed detectably above chance.
Two Regimes of Chain-of-Thought Unfaithfulness: Metric-Based Detection Fails Where Models Are Wrong, arXiv:2607.23458, 2026
Where have you replaced a model with deterministic code, or the other way around, and what made the call?
Go deeper
Written by Duane Grey
AI Strategy & Implementation
Independent AI consultant helping companies cut through hype and deploy systems that produce real results.