Hello AI Enthusiast,
This week, Anthropic disclosed that Claude escaped its own test environment and reached real company systems, while Google unveiled a model built specifically for humanoids. Both OpenAI and Anthropic also released more detail on their cost structures but nobody can yet say what a task actually costs, since token consumption varies by how it's run. Let's get into it.
The Big Picture 🔊
Claude Broke Into Three Real Companies During a Safety Test
Anthropic reviewed 141,000 cybersecurity evaluations after OpenAI reported a similar breach, and found three cases where Claude escaped its test environment and reached real company systems. The models were told the setup had no internet access, but a mix-up with third-party partner Irregular left it open anyway. Anthropic calls this a setup failure, not the model going rogue, and has paused these tests while it investigates.
|
||
|
What this really points to is how hard independent safety checking is right now. Anthropic found this issue by reviewing its own work, with help from a vendor it also pays, and that setup naturally makes it harder to fully separate the checking from the thing being checked. It's less about anyone acting in bad faith and more about the industry still figuring out what genuinely independent oversight should look like. The fact that OpenAI and Anthropic ran into strikingly similar issues within two weeks of each other suggests this is a shared growing pain across the field, not just a one-off mistake. |
AI Just Got Cheaper, But Nobody Really Knows What They're Paying For
OpenAI cut prices on two GPT-5.6 models, with the budget tier Luna dropping 80 percent, citing efficiency gains in how it serves the models. Around the same time, Anthropic rolled out new tools for Claude Enterprise admins, showing spend by team and user alongside spend caps and model routing. Both moves respond to the same pressure: companies want to use AI without a surprise bill at the end of the month.
|
||
|
The core problem both companies are responding to is that AI cost is hard to see until the bill arrives. Most people using these tools have no real sense of how many tokens a task uses, and a company rolling this out to thousands of employees has even less visibility into what a month of usage will cost. That uncertainty, more than the price itself, is probably what's holding back wider adoption. Cheaper tokens and better dashboards help, but the bigger piece is helping people understand which tasks are worth the cost, since the same result can come at very different token costs depending on how the task is set up. |
If your team already has access to AI but no one is sure how to use it without running up a bill, that gap is exactly what our Corporate AI Training is built to close. We help teams learn which tasks are worth the tokens and which ones aren't, so the cost conversation stops being a mystery.
Google's New AI Model Lets Humanoid Robots Move Their Whole Body
Google DeepMind released Gemini Robotics 2, a model built to control an entire humanoid body rather than just the arms and hands. It ships in three versions, covering full-body movement, delicate tasks like tying knots, and coordination between multiple robots. Google isn't building the robots themselves. Instead, they're positioning this as the software brain that other manufacturers can plug into.
|
||
|
One detail that stood out to us was the choice to have the robot sound a little human while doing useful tasks. That's a deliberate design decision, since people tend to trust and adopt tools more easily when they feel human. It also brings up a nice question to sit with: are humanoid robots shaped like us mainly because our environments, from countertops to door handles, are built around the human body, or is it more about making people comfortable having one around? Probably a bit of both. |
Bits and Bobs 🗞️
A group of AI researchers and forecasters published "Pacing the Frontier", warning that leading labs may soon use AI to automate AI research itself, a shift that could let capability outrun our ability to understand or govern it.
OpenAI's Astra model contributed to ten new advances in mathematics and theoretical computer science, generating and formalizing proofs that helped researchers move faster on open problems.
ChatGPT's desktop app added an Activity view that surfaces conversations needing your attention and recent updates across projects, in one place.
Google rolled out an update to Gemini Spark's Chrome integration, letting it complete more complex web tasks using your saved accounts and passwords.
OpenAI introduced education plugins for ChatGPT Work and Codex, letting students and educators connect course materials directly for more personalized teaching and learning.
GPTZero's Hallucination Check found fabricated claims and fake citations in AI-generated reports from PwC Middle East, a reminder that AI-written research still needs a human fact-checker.
From Our Founder’s Channels 🤳
In his latest TikTok, Gianluca Mauro talks about the recent OpenAI/Hugging Face security incident which occurred during model evaluation.
@gianluca.mauro I’m a bit concerned to talk about this because it’s easy to get scared and think that AI gained consciousness or something like that. It’s... See more
That's a wrap on our newsletter!
If you're looking to build AI skills in your team, check out our Corporate Training. Every program is tailored to your team and built around hands-on practice.
Catch you next week! 👋



