The belief I started with

I build construction software. The product, Tride, is used by two construction companies, and most of it exists because I watched site teams run procurement, inventory and approvals through Excel sheets and WhatsApp threads.

One part of it reads documents: invoices, delivery challans, attendance registers. Take a single supplier invoice. Pulling the vendor, line items and totals off the page is extraction. Checking whether the quantities match what the site actually received is a lookup and a comparison. Deciding who has to approve it is a rule. Knowing whether this vendor's invoices usually need correcting is closer to prediction. And every so often an invoice shows up that fits no pattern, and someone has to use judgment.

At Tride, those pieces already behave differently. The AI extracts. When it isn't confident, a person reviews the result, and approved corrections get reused for that vendor next time. Permissions, transactions and approvals don't involve a model at all. They're ordinary code.

So the question I kept coming back to was simple: why should one maximum-capability model own every one of those steps?

My first answer was small language models. Frontier models are the right tool for broad reasoning, coding and genuinely new problems. But a company's recurring work, the same kinds of documents, policies and decisions over and over, seemed like it should eventually move to smaller models trained on the company's own data. They'd be cheaper, easier to run privately, and better at the narrow job.

I spent a lot of time going through the research to see if that held up. Part of it did. But the more I read, the less "small" seemed like the right word for where this work ends up. I still think well-understood work moves away from always using the most capable model. I'm much less sure the destination is a small model.

What held up

The first part held up. A model trained for a narrow task can beat a much more general one at that task.

The cleanest evidence I found is a 2025 EMNLP study by Pecher and colleagues. They compared specialized small models against prompted general-purpose LLMs on text classification. Once the specialized models had enough labeled examples, they often matched or beat the general ones. On average that took around a hundred labels, but the number swung a lot by task, and accounting for that variance pushed it up by 100 to 200 percent. That's classification against prompted baselines, so it doesn't settle small versus large in general. It does show a specialized model doesn't need a huge dataset to catch up.

A more recent example is a May 2026 preprint by Amirhossein Yousefiramandi and Ciarán Cooney at Clarivate, on patent classification. A Llama-3.2-3B model fine-tuned with LoRA scored 0.860 and 0.785 F1 on two patent datasets. The best prompted closed model, Claude Opus 4.6 with five examples, scored 0.823 and 0.752. GPT-5-family models, as routed through Databricks, did worse, and the smallest of them peaked at 0.675 and 0.633.

The caveats matter a lot here. It's an industry preprint on one narrow domain, the comparison is lopsided by design, and the authors report that some fine-tuning runs failed to converge. The fine-tuned model learned from roughly 1,500 to 1,700 labeled examples, while the API models saw five in their prompt. So I wouldn't summarize it as a 3B model beating frontier models. I think a fairer summary is that a small model that had seen the task's real data beat large models that had seen almost none of it.

A January 2026 preprint on enterprise search by Yue Kang and colleagues points at where that advantage comes from. They fine-tuned Phi-3.5-mini to judge how relevant documents were to search queries, and checked it against 923 human-labeled pairs. It ended up on par with GPT-4o, the model that had generated its training labels. Training on public data barely moved it. What moved it was synthetic data seeded from the company's own documents and real query patterns.

That's the part of my original idea that survived. On a bounded task, a model that has seen the task's own distribution can have an edge that general capability doesn't automatically erase.

What broke

This is where my original argument started to break.

Specializing a model makes it better on the inputs it was trained for. It can also make it worse on inputs that look different. Calderon and colleagues measured this across more than 14,000 domain shifts and 21 models (Findings of EMNLP 2024). Fine-tuned models were strongest when the test data looked like the training data. When the domain shifted, few-shot LLMs often held up better. Both kinds of model degraded, so general models aren't immune to shift. They were just less brittle in that study.

That matters for enterprise work because the inputs move. A new vendor sends invoices in a format nobody has seen. A policy changes. A site starts using a different register. A specialist trained on last quarter's documents doesn't know any of that happened.

The second problem is recovery. An ACL 2026 paper on multi-turn tool use by Zhiwei Zhang and colleagues (Fission-GRPO) reported that on one benchmark, BFCL v4 multi-turn, Claude Sonnet 4 recovered from its own earlier execution errors more than 50% of the time, while Qwen3-8B recovered about 20% of the time. The paper's training method improved the 8B model's recovery by 5.7 points, which still left a large gap. That's one benchmark and one pair of models, so I wouldn't stretch it far. My read, and the paper doesn't test this, is that smaller specialized models may be weakest exactly when something goes wrong, or when the input drifts away from what they were trained on.

The third problem hit the part of my idea I was most attached to: the company's own knowledge living inside its own model. In a 2024 EMNLP paper by Ovadia and colleagues, retrieval beat unsupervised fine-tuning at getting facts into a model, though showing the model many rephrasings of the same fact helped fine-tuning. In a controlled 2024 study by Gekhman and colleagues, new facts taught through supervised fine-tuning were learned slowly, and the model's tendency to hallucinate rose as it learned them. And an ICLR 2025 paper by Hu and colleagues found that current approximate methods for making a model "forget" something mostly suppress it: in their tests, fine-tuning on a small, loosely related dataset brought the knowledge back. None of this means facts can't go into weights. It means knowledge that changes, needs a citation, or might have to be deleted is awkward to keep there. Vendor prices change. Approval policies change. Some records have to be removable.

So specialization really does help inside one region of the input space. The catch is that the region you can trust gets narrower, and in enterprise work you can't ignore what falls outside it.

The variable that matters more than model size

After all of this, I stopped asking how small the model could be. The more useful question was: what's the cheapest mechanism that can reliably handle this part of the task?

That question has more answers than "small model." This is the rough framework I use now. It's how I think about it, not a taxonomy from any of the papers.

If the step is an explicit rule or a guarantee, it should be code. Who can approve a purchase order above a certain amount isn't a judgment call, and I don't want a probabilistic system making it.

If the output is a fixed set of labels or a number, a conventional classifier or predictive model is often enough. For predictions on structured tables there's even evidence against reaching for an LLM: a 2026 PNAS Nexus paper by Liu, Yang and Adomavicius found that task-irrelevant changes, like renaming variables, swung LLM prediction error by as much as 82% in some settings.

If the output has to be generated but stays bounded, like turning a messy document into a structured record or taking actions from a fixed set of tools, a specialized generative model can make sense. This is where small language models actually fit.

If the work is open-ended, ambiguous, shifting or hard to check, a general or frontier model is usually the safer default today.

How I now assign AI work

more ambiguity, harder to check the output

more deterministic and easier to check

  • Deterministic rule

    Use code when the logic is an explicit rule or must be guaranteed.

    Examples

    • approval thresholds
    • permissions
  • Predictive / classification model

    Use a trained model when the output is a fixed label or a number.

    Examples

    • fixed labels
    • scores
    • tabular predictions
  • Specialized generative model

    Use a task-specific model when the output must be generated but stays bounded.

    Examples

    • messy document → structured record
    • actions from a fixed tool set
  • General / frontier model

    Use a general model when the task is open-ended, ambiguous, shifting, or hard to check.

    Examples

    • novel cases
    • ambiguous cases
    • shifting cases

more open-ended and harder to check

more ambiguity, harder to check the output

more deterministic and easier to check

more open-ended and harder to check

My framework, not a taxonomy from the research. Only hand a task to something further left if you can check its outputs on real inputs.

For me, the thing that decides how far a task can move toward the cheaper end is checkability. If I want to replace a stronger component with a cheaper one, I need a way to tell whether the cheaper one still works on the real inputs it will see. Without that, moving down is a guess. It's also why, at Tride, uncertain extractions go to human review. The review queue protects the data, and the corrections people approve give the system better examples for that vendor next time.

Size itself turned out to be a slippery word. OpenAI's open-weight gpt-oss-120b has 116.8 billion parameters, but only 5.1 billion are active for any given token, and with its MoE weights quantized to MXFP4, it fits on a single 80GB GPU. Whether that's big or small depends on whether you care about memory, compute per token, or how much the model knows. For most teams calling an API, the size they actually experience is price and latency.

Literal size still matters when a model has to run on a device, offline, in an air-gapped environment, or on hardware you can't change. Outside those cases, parameter count tells you surprisingly little about what a component is for.

The fight I didn't expect

Once I'd moved from "small models" to "specialized mechanisms," I expected the hard part to be specialists versus frontier models. The harder fight turned out to be specialists versus cheap general models.

The strongest version of the other side goes like this. Bounded work does move away from maximum frontier capability, but it moves to the cheap tiers that model vendors already sell. Nobody has to collect labels, train anything, or keep a custom model alive.

OpenAI is clearly betting on this. It positioned GPT-5.4 nano for exactly the work I've been describing: classification, data extraction and ranking. Then on October 1, 2026, it deprecated nano, scheduled its shutdown for April 1, 2027, and named GPT-6 Luna as the replacement, priced at $0.10 per million input tokens and $0.50 per million output tokens, half of nano's input price. That's positioning and pricing, not evidence that cheap models win. But it shows how fast this tier moves. A team building a specialist may find the cheap baseline it was trying to beat has already been replaced by a cheaper one.

At the same time, OpenAI is winding down self-serve fine-tuning. Starting May 7, 2026, organizations that had never run fine-tuning could no longer create fine-tuning jobs, and active existing customers can't create new ones from January 6, 2027. Fine-tuned models keep running only until their base model is deprecated. I don't read that as proof specialization doesn't work, since there are other ways to specialize. I read it as a reminder that a specialist built on someone else's platform can have an expiry date you don't control.

The benchmark evidence from earlier is weaker against this argument than it first looks. In the patent study, the cheapest GPT-5-family model was well behind the fine-tuned model, but it saw five examples. I couldn't find a study that tested a cheap general model given hundreds of retrieved examples from the real task.

Then there's the hole I couldn't fill. I didn't find a strong independent study that compares a modern specialist against a current cheap general tier while counting everything: labels, training, evals, engineering, serving, monitoring, retraining and upgrades. Upgrades aren't free either. A Findings of EMNLP 2024 paper by Echterhoff and colleagues found that when a base model is updated, fine-tuned task adapters start getting some previously correct cases wrong, even when the fine-tuning recipe stays identical.

So the patent result tells me a specialist can win on accuracy. It doesn't tell me whether that specialist is still the better choice after a year of keeping it alive, next to a cheap model that improves every few months without anyone touching it. Accuracy on a benchmark is the easy question. The hard one is whether the specialist stays ahead once you count the whole lifecycle, and I couldn't find a study that has measured it properly.

My bet, and what would change my mind

Right now, here's where I've landed.

I think well-understood, measurable enterprise work will increasingly stop running on maximum frontier capability by default. It will move to the cheapest mechanism that can reliably handle it. Sometimes that's code, sometimes a classifier, sometimes a specialized model, and sometimes a cheap general model with good retrieval. I also think frontier models will stay disproportionately valuable for the parts that are new, ambiguous, shifting, or hard to recover from.

On the open question, I'm betting specialized mechanisms capture a meaningful share of that work, rather than cheap general tiers absorbing nearly all of it. I don't have proof for this, but I do have reasons. A lot of enterprise work runs on distributions specific to one company and repeats at high volume. The outputs can often be checked, and the hard constraints can live in code. Over time a company accumulates real outcomes and corrections that a general model doesn't know by default. At enough volume, small per-call differences in accuracy and cost add up.

Here's what would change my mind. If cheap general models, given enough context and examples from the real task, keep matching specialists on real enterprise distributions, and stay cheaper once the whole lifecycle is counted, most of the case for specialized models goes away. I'd also update if specialists kept degrading under normal business drift faster than teams could realistically retrain them.

I started out asking how small the model could get. The question I'd ask now is how much general intelligence a task still needs once you understand it well enough to check the answer.

Sources

Research papers

Product and platform sources