
How much control are you willing to give an AI model over your digital life? Getting real value from an agent means handing it the keys. For those who hesitate, that can feel risky. Inside OpenAI, the teams building the company’s ChatGPT Work product have already made that leap. The lead engineer on the desktop app has connected the assistant to his inbox, Slack, phone, Notion, Figma and more. He says it is the only way to test the future.
ChatGPT Work, released last month and available on OpenAI’s lowest subscription tier at $20 a month, is the company’s biggest commercial bet for reaching white-collar workers. It is not a chat model that waits for a question. It connects a large language model to the software tools used by accountants, investors, doctors, project managers and others whose days are spent inside browsers and business apps. The model is meant to take on multistep projects, pulling information from connected systems and producing finished outputs.
OpenAI’s product message describes a world where artificial intelligence goes beyond answering questions to help people turn their biggest ideas into reality. The potential payoff for the company is clear: an agent that works for hours, consults many tools and generates thousands of tokens is far more lucrative on a per-user basis than a chatbot that answers a few queries.
From coding assistant to office worker
ChatGPT Work is a modified version of Codex, OpenAI’s coding tool. Software engineers were the first to get comfortable with agents that could write software, run tests and debug code. The challenge is translating that functionality to people who do not write code. Early versions of Codex were actively hostile to non-engineers, according to the app’s lead engineer, showing them empty diffs and asking questions about code. The team spent months making the product more general purpose.
The company sees an enormous market beyond developers. Codex demonstrated that AI can master a job that is digital, measurable and, in the end, a code generation exercise. But law, finance, medicine, operations and marketing involve messy systems, private information and judgment calls. Tools that specialize in those verticals, such as Harvey for law and Clay for sales, are already chasing those customers with a model-agnostic approach. They will use whichever AI model works best at the time. The competition could keep value away from the largest labs if they cannot quickly assemble the complementary assets needed to scale AI in the market, a point made by analysts.
An adoption gap inside and outside OpenAI
Even OpenAI admits that adoption has not yet matched the hype. An OpenAI-backed study found that in June, 98 percent of OpenAI employees were using Codex, but only 17 percent of organizational subscribers and less than 1 percent of individual subscribers had tried the agentic coding tool. The contrast between near-total internal use and negligible external use is the challenge the company is now trying to solve.
OpenAI’s product chief says the more value and utility the company generates for users, the more willing those users will be to pay for it. He points to ChatGPT itself as an example. But enterprise-wide adoption for non-engineers requires painless setup, intuitive controls and an obvious return on effort. None of those are guaranteed.
One internal goal is to reduce the complexity around the “magic box,” the plain chat interface where a user types a request and receives a result. OpenAI has added buttons for projects and plug-ins, but the long-term vision is to let the model figure out what it needs without users managing every detail. The product lead compares this to skeuomorphism, the design practice that made early calculators and note-taking apps look like their physical counterparts. The extra visual cues helped people make the transition from old tools to new ones.
What ChatGPT Work can actually do
In its current form, ChatGPT Work is best at routine, data-intensive coordination. Internal employees use it to prepare weekly metrics reports, turn spreadsheets into planning tools, create charts from Slack conversations about engineering problems, and assemble investment memos. Users can grant the agent access to a Google Calendar, Salesforce, GitHub and dozens of other services. With that access, the agent can draft documents, send emails, update records and perform research across sources.
A journalist testing the product asked it to read a preschool calendar from an inbox and create matching events in Google Calendar. The agent did it, eliminating tedious data entry. The same test produced an auto-updating dashboard of financial metrics for publicly traded companies, built a searchable database of space launches, and began sending a weekly email about new academic AI research. For that to work, the user must be willing to share sensitive information, including private messages and drafts. The engineer who runs the product said he would accept the occasional personal data hit to do the job, though he added that he has not had to do so far.
There are still rough edges. Permission settings are confusing. In one attempt to grant read-only access to a cloud drive, the system returned repeated error messages before a mobile dialog box explained that full access was required. Some settings only appear in the web application, forcing users to switch between devices. The agent can add events to an existing calendar but cannot create a new calendar. The product team also acknowledges that new users do not always know which effort level setting to choose, leading to weak results from the lowest settings.
Learning from Claude’s success
The design of ChatGPT Work was shaped by a competitive stumble. OpenAI built Codex, its coding agent, before Anthropic’s Claude Code appeared. That early version assumed the model was smart enough to complete a whole coding assignment with almost no interaction. The result was often wrong, and it required technical expertise to correct.
Anthropic designed Claude Code to maintain a back-and-forth conversation with the user. It would survey the options, present three or four possible directions, and check in after each step. That reduced mistakes and gave the user a role in steering the work. The approach proved more effective. OpenAI later added more checkpoints and interaction to Codex, and the same philosophy now carries over to ChatGPT Work.
Despite similar interfaces and a prompt that asks new users to import their data from rival products, OpenAI’s engineers deny spending much time studying competitors. They say the obvious difference is model quality. A better model needs less hand-holding and can leverage the harness more effectively.
What makes a good harness
An LLM does not act alone. Software wrapped around it, sometimes called a harness, controls what data the model sees, what tools it can call and how it presents results. A good harness filters out noise and gives the model exactly the context it needs. In the view of one OpenAI engineering lead, the harness is not the long-term moat. The next model will arrive in a few months and make elaborate shortcuts obsolete. Simpler is better.
Outside researchers are less convinced that any single model maker can own the entire experience. An open source harness called Pi, built by the software company Earendil, has been used in major projects and outperformed Codex on some benchmarks while running the same GPT 5.5-level model. Its creator argues that minimalist harnesses can work for coding, and that the real bottleneck is training data. Work tasks in management or operations rarely end in a clean, digital log. If models cannot observe the outcome of a decision made months earlier, they cannot learn to improve.
OpenAI responds with its own evaluation benchmark, GDPval, built from 44 occupations and hundreds of knowledge-work tests, supplemented with feedback from users. The company also watches its own workforce. Product engineers have to distinguish between workflows that are universal and behaviors that are strange because OpenAI is an AI company. Early external users who allow their actions to be recorded will provide far richer evidence, unless they opt out of training.
Money, token burn and subsidies
Cost is hard to hide. During four days of testing on a $20 subscription, one account consumed more than 80 million tokens. The model estimated the usage to be worth about $65, more than three times the monthly subscription price. There is no dashboard in the app to alert users to the subsidy they are receiving.
OpenAI says efficiency will lower the expense over time. It recently cut prices by 80 percent for users of its Luna model. The company predicts that the same tasks will cost less in six months as efficiency improves. But large-scale adoption of agentic AI will force harder conversations about pricing, especially if power users consume hundreds of dollars of compute each month.
The agent market is only beginning. For now, ChatGPT Work works best when tasks are concrete, digital and repetitive. Whether it reaches a billion users depends less on clever interfaces and more on whether ordinary workers are willing to share their inboxes, calendars, files and private messages with a model that acts without asking permission. The engineers who built it are willing to take that risk. Most early users have not yet decided.
Source:TechCrunch News
