What we found running the same six tasks across hosted and open weights for a year.
On narrow, repetitive tasks with a house style — claims triage, document classification, field extraction — a small model tuned on a few thousand of your own labelled examples beat the largest hosted model in five of six cases. Not on raw capability; on consistency, latency and cost per item.
The caveat matters: this holds when the task is narrow and the labels are yours. Widen the task and the ranking flips. Which is why we benchmark against your data before anyone signs a licence.