Google Ships Gemini 4 Argon, but Wall Street Wants a Consumer Agent
Google's Gemini 4 Argon ties OpenAI on a key cybersecurity benchmark, but JPMorgan says the company still lacks a consumer agent to match Meta's Muse and OpenAI's Dots.
Updated
Why it matters
- Google unveiled Gemini 4 Argon on Wednesday, tying OpenAI on a key cybersecurity benchmark and posting leading software engineering results, at $2 per million input tokens and $10 per million output tokens.
- Meta's Muse app topped ChatGPT on Apple's App Store and passed 5 million downloads as of Sept. 30, while Alphabet stock is down about 6% over three months as Meta rose 19%.
- Google's agent Spark remains limited to paying subscribers and can't make outbound calls or complete purchases; Argon has no public release date while it undergoes voluntary U.S. government safety testing.
Google unveiled its long-awaited Gemini 4 Argon model on Wednesday, promising major advances in coding, cybersecurity, and complex tasks — but Wall Street's verdict arrived within a day: a frontier model is no longer enough without a breakout personal agent.
According to industry benchmarks, Argon ties OpenAI on a key cybersecurity test and posts leading results in software engineering. Its introductory pricing of $2 per million input tokens and $10 per million output tokens matches OpenAI's newly discounted GPT-6.1 Sol model. A token is equivalent to about three-quarters of one word.
The release lands at a moment when the market has shifted its attention from raw model capability to what people can actually build with the technology. Personal agents — tools that carry out tasks like scouring emails, booking travel and organizing expenses on a user's behalf — are surging in popularity among businesses and consumers. That shift explains why a technically competitive model launch is meeting a skeptical reception from investors.
The numbers that frame the story
The stock market has already rendered a judgment on the two companies' trajectories. Over the past three months, Alphabet's stock is down about 6%, while Meta is up 19% over the same stretch, sparked by September's rally, which was the sharpest for any month since 2022.
Meta's momentum comes largely from its Muse app, which launched last month and soared to the top of Apple's App Store, topping ChatGPT. As of Sept. 30, Muse had reached more than 5 million downloads, according to Sensor Tower. Earlier this week, OpenAI rolled out Dots, its own personal agent offering, capitalizing on the same demand.
Analysts at JPMorgan Chase wrote in a note on Thursday that Google needs a significant advance in its personal agent offerings to generate consumer enthusiasm. Bank of America identified the potential for Gemini 4 to strengthen Google's cloud business and existing products, while providing a foundation for a future personal agent.
The JPMorgan team was blunt about the gap. "Google has struggled to translate ongoing product velocity into splashy releases that resonate with consumers, capture sustained mindshare, & meaningfully drive sentiment around AI leadership," the analysts wrote.
Google's home-court advantage
In the consumer market, Google has a structural edge that no rival can easily replicate. Billions of people use Google products, including Gmail, Calendar, Chrome, and Search. Their messages and schedules are already stored within Google's ecosystem, potentially giving the company a substantial advantage in developing an assistant capable of working across those services.
Google's agent, Spark, launched in May at the company's I/O developers conference. It can work across Gmail and Calendar, navigate websites through Chrome, and complete tasks such as filling out online forms. It's also available through Google's mobile and desktop apps.
But Spark has two constraints that limit its reach. First, it remains limited to paying subscribers, while Meta's Muse is free with usage caps — a difference Google itself acknowledges. In a background briefing with CNBC, members of Google's agent team outlined Spark's existing capabilities and the company's approach to expanding access, and conceded that Muse's free availability has helped consumers discover and experiment with it.
Second, Spark is functionally narrower than its rival. Unlike Muse, Spark can't make outbound phone calls or complete purchases on a user's behalf. Instead, it takes users through the purchasing process before handing control back to them for final approval — a deliberately conservative design that reflects the sensitivities around granting AI assistants access to sensitive personal information.
Those sensitivities are not hypothetical. Reuters reported in September that Meta had experimented with a human concierge system in which contractors handled some phone calls when Muse couldn't complete them autonomously. Internal privacy concerns prompted Meta to suspend the experiment, while its automated phone-calling feature remains in beta. Meta didn't immediately respond to a request for comment.
Can Argon strengthen Spark?
Google is now evaluating whether its latest frontier model could give Spark new capabilities. In an interview with CNBC, Tulsee Doshi, head of product for Gemini, highlighted Argon's ability to tackle complicated assignments that require multiple steps and can run for extended periods. She said Google is still evaluating where Argon can be deployed most effectively.
The engineering reality is more layered than simply bolting the biggest model onto the agent. Personal agents don't need the most powerful frontier models to complete everyday tasks, and prior to Argon, Google spent much of the summer releasing cheaper, lighter models designed to handle routine requests at scale. Doshi said those models remain essential for everyday agents, while more powerful frontier models can provide the additional reasoning capabilities required for complex assignments.
Then there's the cost side. Running powerful AI models is expensive, particularly for agents that operate continuously and perform multiple tasks. Google's introductory Argon pricing is competitive, but the company said its published rates will eventually double to $4 per million input tokens and $20 per million output tokens. One challenge for Google is determining how to combine its models into an agent that can handle increasingly sophisticated tasks without making the service prohibitively expensive to operate.
Limited availability, limited verification
Argon also isn't broadly available yet. Google is participating in a voluntary U.S. government safety-testing process and initially restricting access to select cybersecurity defenders and enterprise cloud customers. The company plans to expand availability to developers, enterprise customers and consumers, but hasn't announced a public release date.
That restriction is limiting independent testing of Argon's capabilities. Bloomberg reported Wednesday that some Google employees had questioned the model's real-world coding performance despite strong benchmark results. Google disputes that characterization, and told CNBC that employees across the company have been testing versions of Gemini 4 for weeks, with some receiving unlimited access.
The safety-testing posture also connects to a broader debate at the frontier. Attention at the model level has turned to safety in recent weeks, with Anthropic CEO Dario Amodei sparking a firestorm three weeks ago by urging the top AI labs to slow the pace of development as concerns intensify about their potential dangers.
The stakes
The core tension for Google is this: it may be back at the frontier of artificial intelligence on benchmarks, but investors and consumers increasingly care about what people can build with the technology. OpenAI has Dots. Meta has Muse and 5 million-plus downloads. Google has a technically capable agent locked behind a paywall, a frontier model with no public release date, and a stock down 6% while its closest consumer rival is up 19%.
Google's answers to those pressures are now visible in outline — Argon's multi-step reasoning as a potential engine for a more capable Spark, lighter models handling routine requests at scale, and an ecosystem of Gmail, Calendar and Chrome that no competitor can match for reach. Whether the company can combine them into something consumers actually notice, before OpenAI and Meta consolidate their head start, is the question that will determine whether Argon marks a turning point or another strong benchmark result that fails to move the market.
Original: reuters.com
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
170 articles
Related articles
- Google Launches Gemini 4 Argon, Its Most Advanced AI Model Yet
- Google Announces Gemini 4 Argon, Its Next Frontier Model
- Google Launches Gemini 4 Argon, Undercutting GPT-6 Astra on Price
- Google's Gemini 4 Argon Catches GPT-6 Astra, but Claude Opus 5.5 Stays Ahead
- Google Announces Gemini 4 Argon, But No One Outside Can Use It