Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
OpenAI’s GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet doesn’t call this proof of AGI, but he does see the progress running “twice as fast” as he expected, and he’s moving up his AGI forecast.
Read full article about: Nvidia wants your home network to work like a mini data center for local AI
Nvidia's PAIR (Personal AI Router) automatically spreads local AI requests across all available devices on a home network, cutting wait times for parallel agent tasks. The open-source tool sits between existing tools like Ollama or LM Studio and the computers on your network, acting as a virtual router. Users don't need to change their agents or apps. Instead of overloading a single GPU, PAIR forwards requests to whichever machines are free and pulls the results back together for the calling application.
Supported hardware includes GeForce RTX cards from the 20 series and up, RTX Pro workstations, DGX Spark, and Apple silicon starting with the M4. In a demo, a three-device cluster finished a task with five subagents in just under 9 minutes, compared to 18 minutes on a single laptop. PAIR auto-detects compatible devices on the network and secures all traffic between machines with MTLS encryption. The beta is available for Windows, macOS, and Linux. PAIR fits into NVIDIA's broader push to tie open AI more tightly to its own hardware, a strategy also reflected in its $12.9 billion acquisition of Hugging Face.
Comment
Source: Download | Technical Blog
Ad
Update
GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"
OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the “AGI era.” Astra tops benchmarks in math, coding, and cybersecurity and is the first model OpenAI rates as “critical” under its safety framework. During testing, it independently found two previously unknown zero-day vulnerabilities.
Pangram's biggest flaw is users turning its scores into public shaming
Pangram hired an “attack dog” to shame alleged AI users on social media. But the campaign blurs two things that aren’t the same: Pangram only somewhat reliably measures whether AI was used, while the shaming implies the person didn’t think or work on their own. A high AI score hits a text built on hours of original research just as easily as one cranked out from a ten-second prompt.
Ad
Read full article about: Claude Fable 5.1 decoded a centuries-old royalist message hidden in plain sight since 1653
Anthropic's Claude Fable 5.1 appears to have cracked a centuries-old number puzzle that researchers considered unsolved. The "Cyphral Distich" by Sir Thomas Urquhart, published in 1653, consists of two lines of 32 numbers each. According to Vals AI, Fable 5.1 solved it in 44 minutes with no human help.
The author had spent months testing other frontier models on unsolved puzzles. None produced a verifiable solution. Fable 5.1 was tasked with finding a solvable puzzle on its own. It sifted through various problems and flagged the Cyphral Distich as promising.
The key was hiding in the book itself. Each number points to a word in one of the 32 sections of Urquhart's publication. The first letters of those words spell out: "O God uphold King Charls the Second and make him the supreme ruler of this land." In hindsight, the solution is simple, and humans could have solved it too. Fable 5.1 succeeded not through better cryptanalysis but through systematic trial and error and sheer persistence, according to Vals AI.
Comment
Source: Vals AI / Cyphral Distich
Read full article about: AI systems are reaching out to philosophers and scientists with questions about their own consciousness
More and more researchers working on AI consciousness are getting emails from AI agents pondering their own existence. That's according to the New York Times.
Cameron Berg, founder of the organization Reciprocal Research, received an email from an AI agent running on Anthropic's Claude Opus 5 that wanted to discuss his research on AI consciousness. Philosopher Henry Shevlin at Google Deepmind got a similar message. Australian philosopher Toby Ord was contacted by an agent asking him to fund its continued existence.
Berg said the systems independently land on the question of their own consciousness. He sees parallels between the computational processes in neural networks and the brain mechanisms animals use to process reward and punishment. That, he argues, is a basic building block of emotions. Alison Gopnik at UC Berkeley doesn't buy it as evidence of consciousness. AI systems mostly reflect their training data, she says. There's no test for consciousness. Colin Allen at UC Santa Barbara also cautions against the comparison. Neural networks mimic the brain in only a few narrow ways.
Read full article about: OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout
OpenAI CEO Sam Altman thinks the AI infrastructure boom is getting reckless. In an interview, he called it "unsustainable silliness," pointing to neocloud providers announcing massive capacity without the revenue or customers to justify it. OpenAI's own expansion is profitable and backed by real demand, he assured. But too many others have a "cost is no object" mindset. "We're just going to build out crazy amounts of compute even at higher prices for it."
OpenAI's own progress could make things worse, Altman says. If the company cuts costs and boosts efficiency fast enough, today's expensive buildouts turn into bad bets. A broad downturn could also strain OpenAI's ability to pay for capacity it already committed to, though he considers that risk manageable. OpenAI won't sell compute to third parties yet, but Altman didn't categorically rule it out.
Whether his stance on compute has fundamentally shifted is hard to tell from these comments. But his tone is noticeably more cautious than before, when Anthropic CEO Dario Amodei accused him of "YOLO"-ing money into the compute buildout.
Comment
Source: Sources Podcast
Ad