ATT Open Model Cost Cut
AT&T quietly proved something most enterprise AI budgets are built to ignore.
The company cut its internal AI coding costs by 56% by routing routine queries away from frontier models like Claude and GPT-5 and toward open-source models such as Llama, Nemotron, and Gemma. The performance hit was 2%. The platform processes roughly 45 billion tokens a day.
Today about 40% of employee queries go to open models. AT&T's target is 60 to 70%.
This is not a story about open models being "good enough." It is a story about architecture beating model choice. Most enterprises treat model selection as a single decision made once, at the top, for every use case. AT&T treats it as a routing problem, solved per query, in real time.
The implication for procurement is bigger than the cost line. If routing infrastructure becomes standard, the frontier labs lose pricing power on the largest, lowest-value share of enterprise volume, the routine, repetitive queries. Frontier models keep the hard problems. Open models absorb the rest.
Every company paying a flat per-seat or per-token rate to a single vendor right now is, in effect, using a scalpel to cut lawns.
The real question is not which model is best. It is whether your organization has the routing layer to make that question irrelevant for 60% of its traffic.