Ad
Skip to content
Read full article about: OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity

OpenAI employee Vaibhav Srivastav explains when each of GPT-5.6 Sol's five reasoning levels fits. "Light" and "Low" are for quick, clear-cut tasks. "Medium" works for planning and analysis. "High" and "xhigh" handle complex, multi-step work or "careful verification."

"Max" and "Ultra" work differently: "Max" lets a model spend more time on a single problem. "Ultra" deploys multiple sub-agents in parallel, each tackling a different part of a task. Higher levels take more time and burn through more tokens. Srivastav recommends starting low and only scaling up when needed. The levels don't map to GPT-5.5's tiers, Srivastav says, and anyone switching over should start one level lower than they're used to.

None of this brings OpenAI any closer to its stated goal of making ChatGPT so simple that "almost no interface" is needed. On top of that, Sol's Pro tiers are still missing. Those leaked earlier in a genomics benchmark paper. Even ambitious users will struggle to pick the right level without running their own benchmarks, though the setup may help OpenAI collect usage data.

Read full article about: Tencent moves to buy majority stake in Manus after Beijing forced Meta to unwind its $2 billion deal

Chinese tech giant Tencent is in talks to acquire a majority stake in AI agent startup Manus, according to the Financial Times, after Beijing forced Meta to unwind its $2 billion acquisition of the company. Tencent sees overlap with its own AI agent strategy, including plans to embed an agent into WeChat.

Most earlier investors, including Tencent, ZhenFund, and HSG, plus the management team, are discussing a deal at the same $2 billion valuation. U.S. firm Benchmark is not expected to take part. Manus will keep operating independently out of Singapore and most recently reported annual revenue of close to $500 million.

China blocked Meta's Manus acquisition in April, calling it a violation of investment rules, and imposed an exit ban on founder Xiao Hong. Officials described the deal as a "conspiratorial" attempt to undermine China's tech base and banned foreign investment in Manus. The decision fits into a broader AI arms race between the two countries, where the technology is already being compared to "cyber-nuclear weapons of the AI era" given recent advances in AI-driven cybersecurity attacks.

Read full article about: OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT

OpenAI is already killing its AI browser Atlas, launched just last October 2025. Its features are moving into an updated Chrome extension that lets users run ChatGPT directly in Chrome's sidebar. The company says it's folding in what it learned from Atlas and user feedback. Atlas users will get notified about the switch. Separately, the new desktop "Computer Use" feature lets ChatGPT handle tasks in the background. It can click, type, move files, and work across apps and browsers, either as a one-off action or a recurring task.

When Atlas launched, it looked like a shot at Chrome. Now OpenAI is pulling the plug less than eight months later. That puts the browser on a growing list of scrapped or not very successful OpenAI products, alongside pluginsapps, the ChatGPT Agent, and the Sora video model. For users, bundling everything into ChatGPT might actually be more convenient. But it also means OpenAI has no way to pull users away from Chrome, giving Google a competitive edge thanks to all the browsing data it collects.

Read full article about: Bun ditches Zig for Rust with help from Claude Fable 5, writes over a million lines of code in 11 days

The JavaScript tool Bun has been fully rewritten from Zig to Rust, and Anthropic's Fable 5 did most of the work. Developer Jarred Sumner says the switch came down to reliability. Zig kept producing memory errors and crashes that were hard to fix for good. Rust catches many of those bugs at compile time.

Sumner used a pre-release version of Claude Fable 5 for the project. About 64 Claude instances ran in parallel for 11 days, writing over a million lines of code. The API bill came to roughly $165,000. A human team would have needed about a year, according to Sumner. The new version, Bun v1.4.0, is available as a canary release. It fixes 128 bugs and runs about 2 to 5 percent faster. Sumner didn't have to worry about the hefty price tag. Bun and his team were acquired by Anthropic in December 2025.

Comment Source: Bun
Read full article about: Anthropic's fix for Fable 5's high cost is turning it into a manager that delegates to Sonnet 5

Claude Fable 5 is expensive. Anthropic now recommends using it mainly as a planner, handing execution off to smaller models. The Claude developer team outlines two strategies. In the "Advisor" pattern, Sonnet 5 runs as the executor and only calls Fable 5 when it needs guidance. On SWE-bench Pro, this combo reaches about 92 percent of Fable 5's solo performance at 63 percent of the cost, according to Anthropic. Fable 5 gets called roughly once per task.

Advisor pattern: Sonnet 5 does the work and only consults Fable 5 when needed. | Image: Anthropic

In the second pattern, Fable 5 acts as a planner that delegates tasks to Sonnet 5 worker agents. On BrowseComp, this delivers 96 percent of Fable 5's performance at 46 percent of the cost. Both patterns run through Claude Managed Agents, with each sub-agent using its own cache to avoid duplicate context costs. The official documentation has more details.

Orchestrator pattern: Fable 5 plans and distributes tasks to multiple Sonnet 5 workers. | Image: Anthropic

Anthropic is likely sharing these tips because of growing price pressure. Chinese open-source models are already undercutting Western pricing, and the new GPT-5.6 Sol is much cheaper per token and reportedly more token-efficient too.

Read full article about: Google Deepmind adds background execution and MCP support to Gemini API managed agents

Google Deepmind is adding four new features to Managed Agents in the Gemini API. Developers can now run agents asynchronously in the background using Background Execution, with no open HTTP connection required. Remote MCP (Model Context Protocol) servers can also be connected directly to internal databases or APIs. Another addition lets developers use custom functions alongside the built-in sandbox tools. Finally, credentials like tokens can be refreshed between interactions without losing the sandbox state.

All features are available through the Gemini Interactions API. Code examples for JavaScript, Python, and cURL are in the documentation.

Read full article about: Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year

Chinese AI developer MiniMax is working on a new large language model with 2.7 trillion parameters. MiniMax plans to release the model as open source. That's according to The Information, citing two people familiar with the plans. The model would be larger than any other Chinese AI model currently on the market.

Internally, MiniMax calls it M3 Pro, though the name could change before launch. The two sources say a release could come as early as Q3. The company's current top model, M3, has 428 billion parameters. Larger models tend to perform better on tasks that require complex reasoning and multi-step instructions.

Chinese open-source models have gained traction with developers this year, especially those looking for cheap models to handle high-volume, less critical tasks. MiniMax competes with Zhipu, DeepSeek, and Moonshot AI. Recent reports, however, suggest the Chinese government wants to tighten controls on future releases of such models.

Read full article about: Meta tests always-on AI glasses that capture your entire day

Meta is prototyping AI-powered glasses with a feature called "Super Sensing" that continuously records the wearer's surroundings using cameras and microphones. The glasses constantly capture audio and snap photos every few seconds, according to multiple people familiar with the project. Users could then ask an AI to recall anything they saw or heard, the Financial Times reports.

The project is already sparking internal debate over privacy. Unlike Meta's current Ray-Ban smart glasses, the Super Sensing mode wouldn't activate the LED indicator light. That means bystanders would have no way of knowing when they're being filmed. These plans could still change. Meta is also considering using the collected data to train its own AI models.

Meta previewed related features like "Live AI" at Connect 2025. The idea is for the glasses to build up context throughout the day, factor in earlier information, and help users with tasks. Meta's "Project Aria" research program has been collecting first-person data for AI systems for years, following a similar approach.

The company declined to comment to the FT on internal prototypes but pointed to its privacy-focused technology.

Comment Source: FT
Read full article about: Cohere Transcribe Arabic is an open-source model built for Arabic's toughest transcription problems

Cohere has released Cohere Transcribe Arabic, an open-source model built for Arabic speech recognition. The 2-billion-parameter ASR model is, according to Cohere, the most accurate open-source Arabic speech-to-text system available. It targets the specific challenges of Arabic speech, including dialect variety, bilingual Arabic-English conversations, code-switching, and specialized vocabulary. Cohere says it outscores Whisper Large V3, the standard Cohere Transcribe model, and other systems in benchmarks.

Menschliche Bewertungen arabischer Transkripte auf einer Skala von 1 bis 5: Cohere Transcribe Arabic übertrifft Whisper Large V3 und das Standard-Modell Cohere Transcribe bei Gesamtqualität, Dialekttreue und Code-Switching. | Bild: Cohere
Human ratings of Arabic transcripts on a scale of 1 to 5: Cohere Transcribe Arabic outperforms Whisper Large V3 and the standard Cohere Transcribe model in overall quality, dialect faithfulness, and code-switching. | Image: Cohere

The model ships under the Apache 2.0 license and is available on Hugging Face and through the Cohere API. More benchmarks and examples are on the Cohere blog.

Read full article about: Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens

Chinese AI models are gaining ground with US companies because they cost far less than systems from OpenAI and Anthropic, CNBC reports. Models from companies like DeepSeek and Z.ai are seen as competitive, even as per-token pricing from US providers keeps climbing.

On the platform OpenRouter, Chinese models have accounted for over 30 percent of traffic every week since February 8, hitting 46 percent at times. Last year, the average was just 11 percent. According to OpenRouter employee Justin Summerville, Chinese open-source models run 60 to 90 percent cheaper. The startup Lindy shifted all of its traffic from Anthropic's Claude to DeepSeek. CEO Flo Crivello said the switch saves millions.

Kyle Chan at the Brookings Institution puts the gap between Chinese and US models at six to nine months. That lines up with an estimate from the Center for AI Standards and Innovation (CAISI). The agency published a report in May finding that Chinese AI models trail leading US models by about eight months. The assessment covered cybersecurity, software development, math, science, and abstract reasoning.

Comment Source: CNBC