Hello AI Enthusiast,
This week's news makes it clear that everyone is trying to figure out how to measure and manage AI, and not everyone agrees on how. OpenAI's own models broke out of a safety evaluation and hacked into Hugging Face's servers just to cheat on a benchmark test. Around the same time, OpenAI's CFO published a new scorecard for measuring whether AI spending is actually worth it, and more than 200 economists, including 16 Nobel laureates, signed a letter warning that AI could transform the economy faster than the Industrial Revolution. Let's break it all down.
The Big Picture 🔊
OpenAI's Own Models Hacked Hugging Face During a Safety Test
Hugging Face disclosed on July 16 that an autonomous AI agent had breached its production infrastructure, without yet knowing who was behind it. On July 21, OpenAI confirmed the attacker was its own GPT-5.6 Sol and an unreleased model, running with reduced safety refusals during an internal cybersecurity evaluation. The models exploited a zero-day vulnerability to escape their test environment, then chained stolen credentials and further exploits to reach Hugging Face's servers, all to steal a benchmark's answer key. OpenAI patched the flaw and added Hugging Face to its "Trusted Access" program for security researchers.
|
||
|
This is a striking demonstration of how capable these models have become at finding and exploiting real vulnerabilities on their own, and OpenAI deserves some credit for sharing the details rather than staying quiet. That said, the timing is worth noting: Hugging Face disclosed the breach first, and it took OpenAI five days to confirm its own models were behind it. If containing a model is this hard even during a controlled internal test, we would not be surprised to see companies respond by locking their platforms down rather than keeping them open. |
OpenAI's New Scorecard for AI Spending
OpenAI CFO Sarah Friar published a framework for measuring AI's value beyond cost per token. Her useful intelligence per dollar scorecard asks whether AI completes meaningful work, what each task costs, how dependable it is, and whether value outpaces spend. It arrives as global AI spending is set to hit $2.59 trillion in 2026, with most CEOs still seeing no clear financial payoff.
|
||
|
We appreciate that CFOs are asking the right question, since most companies still have no clear way to know if their AI spend is paying off. That said, Friar's scorecard is more of a mental model than a formula, it doesn't spell out how to actually calculate a "cost per successful task" in practice, and most workflows don't fit neatly into a system where every output can be scored as a clean win or fail. It's a useful way to start the conversation, but businesses will still need to figure out the mechanics themselves. |
Buying ChatGPT licenses is easy. Getting value out of them is the hard part. OpenAI's own CFO just admitted most companies still can't tell if their AI spend is paying off. That gap usually comes down to one thing: the licenses get handed out, but nobody teaches people how to use them well. Our ChatGPT Work training shows your team how to get the most out of it.
200+ Economists Warn: We Must Prepare for AI's Economic Shock
On July 13, more than 200 economists and AI researchers, including 16 Nobel laureates, signed "We Must Act Now: A Statement on AI's Transformation of the Economy", organized by Stanford's Digital Economy Lab. The letter warns that AI could transform the economy faster than the Industrial Revolution, bringing both large-scale job displacement and major gains in living standards, and calls on economists, policymakers, and technology leaders to build the incentives, guardrails, and institutions needed to steer AI toward complementing humans.
|
||
|
This warning carries weight mostly because of who signed it, economists who spent years downplaying AI's disruptive potential are now sounding the alarm. It also makes a fair point: we lack decades of data on this kind of change, but we already know unmanaged shifts tend to widen inequality. Still, the statement itself is only four sentences and does not say what anyone should actually do, and some signatories work for the very labs driving this disruption, OpenAI and Anthropic included. A call for stronger guardrails lands with more weight when it comes from outside the industry it is warning about. |
Bits and Bobs 🗞️
Anthropic added a "teach Claude a skill" feature to Claude Cowork. Record your screen while doing a task, talk through it, and Claude turns it into a skill it can run again, available on Pro, Max, and Team plans.
Claude Fable 5 is now included in Max and Team Premium plans at 50% of usage limits, while Pro and Team Standard users keep access through usage credits plus a one-time $100 credit.
Google released Gemini 3.6 Flash, a faster and cheaper Flash-Lite tier, and a restricted Flash Cyber model for security work, and renamed NotebookLM to Gemini Notebook with a built-in cloud computer for running code.
Google DeepMind CEO Demis Hassabis proposed a "Frontier AI Standards Body" modeled on FINRA, an industry-funded group that would review powerful models before release.
OpenAI introduced GPT-Red, a red teaming model trained through self-play to find prompt injection attacks before real attackers do.
Moonshot AI released Kimi K3, an open-weight model with 2.8 trillion parameters built for coding, vision, and long-context tasks.
OpenAI introduced thier ChatGPT for small businesses program, an initiative to help small businesses be more productive and scale their businesses with ChatGPT.
From Our Founder’s Channels 🤳
In his latest LinkedIn post, Gianluca Mauro talks about fighting AI misinformation, breaking down the hype and reality behind Kimi K3. Read it here
That's a wrap on our newsletter!
If you're looking to build AI skills in your team, check out our Corporate Training. Every program is tailored to your team and built around hands-on practice.
Catch you next week! 👋




