← All signals

The Biggest Change in AI Is Happening Behind Your Chat Window

OpenAI built a chip, Cisco packaged complete AI systems and Google put agents inside real jobs. Here’s how that could change the AI you use every day.

The Biggest Change in AI Is Happening Behind Your Chat Window

Most people experience AI through a blank box and a blinking cursor.

You type a question. ChatGPT, Gemini or another assistant produces an answer. That makes it easy to think that the model is the product and that the company with the smartest model will eventually win.

Three announcements made on August 25 suggest something else. OpenAI published the first performance results from its own AI chip. Cisco expanded its infrastructure with complete, ready-to-install AI racks. Google introduced specialised AI agents for lawyers and financial professionals.

These may sound like separate technology stories. They are really three parts of the same shift.

The next generation of AI will not be determined by the model alone. It will depend on everything beneath it, around it and connected to it.

The chip beneath every answer

OpenAI unveiled its first custom AI chip, Jalapeño, in June. The new information is that the company has now published its first detailed performance results.

Jalapeño is designed for inference. If training is the process of building an AI model, inference is what happens every time that model answers a question, analyses a document or completes a task.

This distinction matters. Training attracts the spectacular investment headlines, but inference happens continuously once millions of people start using the finished product. Every answer consumes computing capacity, electricity and time.

OpenAI tested Jalapeño with three public models: GPT-OSS 120B, DeepSeek R1 and Kimi K2.5. According to the company, its chip completed between 1.5 and 1.9 times more AI work per watt than the Nvidia systems used for comparison. OpenAI also reported between 1.7 and 3.6 times lower end-to-end latency.

Those are OpenAI’s own results, not independent confirmation. The tests use InferenceX, a public benchmarking system, but Jalapeño has not yet proved itself through large-scale commercial deployment.

The consumer relevance is still clear. Faster inference can mean quicker answers and more responsive AI agents. Lower energy use can help providers serve more users with the same power and infrastructure. It does not automatically mean cheaper subscriptions, but it gives OpenAI room to reduce costs or let its agents perform more work.

OpenAI plans to start deploying Jalapeño by the end of 2026. A second generation is already in development, with a third being planned. The company will, however, continue buying and using chips from Nvidia and other partners. Jalapeño creates another option; it does not immediately remove OpenAI’s dependence on external suppliers.

AI is also a physical product

Cisco’s announcement exposes another side of AI that consumers rarely see.

AI does not live somewhere inside an invisible cloud. It runs inside buildings filled with processors, networks, cooling systems and power equipment. All those components need to work together.

Cisco already offered its Secure AI Factory with Nvidia. It is now adding complete rack-scale systems from Supermicro, including dense GPU servers and air or liquid cooling. Cisco combines those systems with its own networking, security, monitoring and management technology.

Customers will be able to buy the Supermicro systems through Cisco from October 2026. The aim is to give businesses a tested package instead of forcing them to assemble an AI data centre from separate components. Cisco says the architecture can support everything from extremely large training projects to inference closer to where data is produced.

This matters because a brilliant model is of little use when a company cannot install, cool, secure or operate the infrastructure required to run it.

Google moves AI inside the job

Google is attacking the same problem from the opposite direction.

Instead of starting with hardware, it is building everything that must sit around a model before a company can trust it with real work.

Gemini Enterprise for Legal connects AI agents to systems used by lawyers, including Microsoft 365, Google Workspace, iManage, NetDocuments, Docusign and Everlaw. The agents can assist with contract reviews, legal research, privacy requests and regulatory monitoring.

Gemini Enterprise for Financial Services connects to financial sources such as FactSet, Moody’s, MSCI, PitchBook and SEC filings. Its Financial Research agent includes more than 50 specialised skills and can produce research with source citations, stated methods and data snapshots for auditing.

Both products are currently in preview. Google has not published broad customer results, independent accuracy tests or public pricing. It does say that customer data, internal business rules, custom agents and their outputs will not be used to train its foundation models.

What this means for the AI you will use

OpenAI is moving down from the model into the chip. Cisco is moving up from the network into the complete AI system. Google is moving beyond the model into professional data and workflows.

They are all trying to control more of the journey between your question and a useful result.

For consumers, the most visible improvements may therefore stop arriving as dramatic new chatbot launches. AI could simply become faster, appear inside more services and complete longer tasks with less waiting.

For businesses, intelligence alone is no longer enough. AI also needs access to the right information, permission to perform specific actions, reliable infrastructure and a clear record of what it has done.

The next AI battle is not only about who builds the smartest brain.

It is about who builds the best system around it.