Itβs Complicated: My First Project With GPT-5.6 Sol on Max Thinking
GPT-5.6 Sol turned a small utility into an enterprise-grade planning exercise, then reminded me that sophisticated agents still miss things.
GPT-5.6 Sol turned a small utility into an enterprise-grade planning exercise, then reminded me that sophisticated agents still miss things.
DentaQuest is notifying more than 15 million people after a three-day intrusion in May. The dataset ShinyHunters published covers 2.6 million.
Anthropic's Opus 5 costs the same as Opus 4.8 and half of Fable 5 β and it is the fifth frontier model to ship in a little over two weeks.
Google DeepMind releases Gemini 3.6 Flash, a high-efficiency workhorse model optimized for production AI agents, coding, and multi-step reasoning.
A bipartisan House bill would make frontier AI developers keep a working off switch and let DHS order it thrown. Two incidents this summer explain why.
The White House accuses Moonshot of distilling Anthropic's Fable to build Kimi K3. Who's claiming what, on what evidence β and what's still unproven.
A security researcher pointed GPT-5.6 Sol Ultra at the WordPress codebase and got back a pre-authentication remote-code-execution chain affecting 500 million sites β for about $25 in subscription tokens. The patch is out, exploitation has started, and the economics of vulnerability research just changed.
Moonshot AI paused new Kimi K3 subscriptions after launch demand hit its GPU ceiling β hours before Axios reported the Trump administration is weighing restrictions on advanced Chinese models. The capacity crunch, the market reaction, and the policy fight, up to the moment.
Anthropic is giving verified US K-12 teachers a free year of premium Claude, with standards-aligned lesson planning and agentic tools included. The pedagogy plumbing is the pitch; the district-governance questions are the caveat.
Moonshot AI released Kimi K3 β 2.8 trillion parameters, native vision, and a 1M-token context window, live in the API now with open weights promised by July 27. Benchmarks, pricing, the self-hosting reality, and what still needs independent verification.