Big tech companies are all building office Agents, but what problems do they actually solve?

ChatGPT_Image_2026%E5%B9%B48%E6%9C%886%E6%97%A5_21_18_03.png

In recent months, domestic office Agents have suddenly become a hot topic.

Tencent has WorkBuddy, Alibaba has QoderWork, ByteDance launched TRAE Work, and even Kimi came up with Kimi Work. These products come from different companies, use different models, and connect to different ecosystems. But if you cover up their names, it's almost impossible to tell which is which.

They all share the same interaction pattern: the user opens a dialog box and states the result they want. The Agent automatically breaks down the task, searches for information, reads files, calls on browsers and office software, and finally delivers a report, spreadsheet, or PPT.

These big tech companies seem to be waging a fierce office AI war, but I've never quite understood what exactly they are fighting over.


What is the essence of an office Agent?

The reason I have this question is that I have been using Codex frequently at a high volume.

From a product structure perspective, these domestic office Agents have not created a way of working that differs from Codex.

Codex handles code repositories; office Agents handle local files, web pages, documents, and spreadsheets. The objects of work have changed, but the fundamental structure has not. In both cases, the user sets a goal, and the Agent executes it.

The differences among the various products mainly come from the models, the ecosystems they can connect to, and the number of role cards and Skills. One company releases a market research expert today, another launches a PPT-making expert tomorrow, and the day after that, a batch of writing, search, spreadsheet, and data analysis Skills go live.

The capability lists keep getting longer, but they are only expanding horizontally.

They have proven they can generate more types of files, but they have not proven they can accomplish deeper, more complex work.

What these products are still best at are tasks with short paths and results that are easy to manually check. For example, organizing a few pieces of material, summarizing a batch of files, generating a report based on existing content, turning a document into a PPT, or writing a few hundred words of daily or weekly reports.

So they are more like Paperwork production tools wrapped in Agent form. By stitching together various file information, they ultimately produce a "report" that isn't too ugly to look at. But these reports neither support judgment nor drive any decision-making.

Therefore, the real homogenization of these products is not just that they all have a dialog box.

None of them have solved a specific problem that requires bearing responsibility for the outcome.

When a product cannot prove its value by results, the first thing it can create is merely a new scenario for consuming electricity.


Why have all products ended up looking the same?

Because this wave of office Agent hype did not originally start from a clear office need. It is an extension of Coding Agents.

Coding Agents have already proven that users are willing to pay for "AI directly executing tasks." Moreover, compared to chatbots, the value of Coding Agents is easier to understand. They can modify code, fix bugs, run tests, and finally deliver verifiable results.

Once this model proved viable, migrating it to the adjacent office market became almost the inevitable choice for every big tech company.

Swap the code repository for local files, swap the IDE and terminal for browsers and office software, swap code for reports, spreadsheets, and PPTs, and you have a general-purpose office Agent.

But replicating the product form is easy; the solution itself cannot be transferred. The business that works for Coding Agents may not necessarily succeed in office scenarios. After all, the concept of "office" is far too broad. It is not a specific scenario; it encompasses sales, finance, product, operations, R&D, legal, and so on.

To solve a specific office problem, one must confront the specific problems of specific users.

For example, an Agent truly built for sales needs to be designed around customers, opportunities, contracts, and payment collection; an Agent for procurement needs to handle suppliers, prices, inventory, and approvals. These segmented scenarios not only have their own processes, documents, and data formats, but also contain evaluation systems with varying standards and their corresponding SaaS tools.

If every office scenario is stuffed into a single Agent, the product becomes extremely bloated and difficult to maintain. So they all chose the same interaction approach: using role cards and Skills to cover as many business scenarios as possible.

This is a "lazy" approach: it does not define outcomes, only the most generic execution, reducing all business problems to document problems.

They copied the shell of Codex, but not the core of Codex, where the model and tools are deeply integrated.

In the end, they all became castrated versions of Codex.


What problems have they actually solved?

For short-path tasks such as organizing materials, processing files, converting formats, and generating content, these office Agents can indeed save some time.

But this portion of value is not enough to explain why Tencent, Alibaba, and ByteDance would simultaneously invest massive resources to develop a batch of products that are highly similar in form and capability.

However, if you look at it from the vendors' perspective, they are actually solving another problem:

How to turn massive AI investments into visible commercial results.

Ordinary AI chat products can only do one question, one answer, while an Agent expands a single request into a complete chain of calls, and most of the time it launches multiple sub-Agents working simultaneously.

Even if the user ends up with just an ordinary report, a large number of model calls have already been generated behind the scenes.

These general-purpose office Agents do not define specific outcomes they are responsible for, nor can they easily prove how much value a report or a PPT is actually worth. So, Token consumption is their most essential proof of value.

Over the past few years, big tech companies have invested heavily in model training, GPUs, data centers, and talent. They must continually prove to internal leadership and the capital markets that AI is not just burning money; it is entering real work scenarios and generating massive model calls, cloud revenue, and enterprise procurement contracts.

Office Agents conveniently provide them with a complete storyline: they generate more model and cloud resource consumption, and that consumption is then used to prove that previous AI investments are converting into usage, revenue, and growth.

So, selling Tokens and sustaining the capital narrative are actually two sides of the same coin.

But the problem is that selling Tokens is like selling electricity. Electricity from the Bay Area is no nobler than electricity from the Three Gorges Dam. Tokens only have value once they solve a specific problem. An Agent consuming more Tokens does not mean it has completed more valuable work.

If an AI tool has no corresponding user scenario, does not solve a specific problem, and cannot define an acceptable outcome, then what it creates is merely an electricity-consuming scenario that serves the capital narrative.

Let's talk

Tell me what you think