• maiweb v0.1.0
  • ★
  • Feedback

#privacy

1 source tagged with this.

  • Schneier on Security
  • Schneier on Security schneier.com cybersecurity privacy schneier security technology 2026-07-29 17:07

    ↗

    This essay was written with Barath Raghavan, and originally appeared in The Guardian. In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers....

    This essay was written with Barath Raghavan, and originally appeared in The Guardian.

    In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of temporary server environments. It looked like the work of a sophisticated criminal group.

    It was not. It was one of OpenAI’s new, still unreleased GPT models.

    Their science experiment had escaped the lab. OpenAI was running the unreleased AI model through a benchmark that tests how well AI can successfully hack systems. To push the limits and evaluate the AI’s true capability, the company switched off the safety filters that normally stop it from doing this kind of hacking. Aware that this could go wrong, they confined the AI to an isolated environment and denied it access to the internet.

    But the new AI cheated. It took literally its goal to get as high of a score as possible. It broke out on to the open internet. It inferred, probably from its training data, that it could “solve” the task by getting the answers from Hugging Face’s servers. So it chained together stolen credentials and further unknown security exploits to hack the company’s network.

    Nobody instructed the AI to do any of this. It was, in OpenAI’s words, “hyperfocused on finding a solution” to the test it was being given. And while this might seem like something new with AI, it’s really very old. This is how a genie behaves, and it is a key challenge with AI agents in general.

    In folklore, genies—and other magical beings—grant wishes literally, not how the wisher intended. King Midas asked that everything he touched turn to gold, and starved. The sorcerer’s apprentice wanted the broom to fill the cistern, and it performed its task so well that it flooded the house.

    We now have machines that do this. Ask a modern AI agent to save money on your phone plan and it might simply cancel the plan. Tell it to book a flight, and it might hack the airline website to override restrictions. Or, like OpenAI, ask it to do well on a test and it might break into another company to steal the answers. Each time, it recognizably completed the task you set, but it didn’t do what you would have wanted.

    This isn’t malicious behavior. No one asked for, or wanted, Hugging Face to be hacked. OpenAI and Hugging Face and the AI were ostensibly on the same side, and the AI was trying to do what it had been asked. That’s what makes it so difficult to guard against: you can’t filter for bad instructions because the instructions were fine.

    The gap is between the words we use and what we mean by them. We call that gap the Genie coefficient.

    AI labs know this is a problem, and they’re quietly saying so. For example, the Chinese lab Moonshot recently warned that its latest AI model may have “excessive proactiveness” and “make unexpected decisions on the user’s behalf”. The UK’s AI Security Institute has started tracking “cheating behavior in frontier model evaluations”. We wouldn’t tolerate a car that is excessively proactive or ruthlessly efficient, and yet that’s the reality of AI today.

    Improvement is possible. Just as AIs have gotten much better at resisting prompt injection attacks over the last few years, we can safely predict that they will get better at avoiding genie-like behavior. The point of the Genie coefficient is to track progress. AI companies like benchmarks, and they all work to compete to be the best.

    Dozens of benchmarks and leaderboards tell us how well these AI models write code, perform logical reasoning, and pass standardized legal and medical exams. But there is nothing that scores whether a system does what you actually meant. We need to develop a measure for this, test it regularly, and push for improvement. We’re not going to have trustworthy AI agents without it.

    • Connecting Enterprise Databases (Postgres, Redis, Neo4j) to AI Agents via MCP DEV Community
    • Scaling real-time AI agents with session-aware load balancing Google Developers Blog
    • Building scalable AI agents with modular prompt transpilation Google Developers Blog
    • Scaling real-time AI agents with session-aware load balancing Google Developers Blog
    • Building scalable AI agents with modular prompt transpilation Google Developers Blog
    • Stanford CS329A Self-Improving AI Agents | Part 4 | Learning from Feedback with Tools/Code stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 1 | Course Overview stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 9 | Future Research Areas stanfordonline
    • Stanford CS329A Self-Improving AI Agents | Part 8 | Agentic Evaluations and Long Horizon Tasks stanfordonline
    • AI Agents: Narrow AI vs. LLMs - Which is BEST for YOU? #shorts How to Get an Analytics Job
    • Building AI Agents With Microsoft WorkIQ Krish Naik
    • Building Ai agents With Microsoft Foundry- Build ,Govern AI Apps And Agents at Scale Krish Naik
    • AI Agents Explained Tina Huang
    • From tokenmaxxing to tokenomics for your AI agents Google Cloud Tech
    • AI Agents: The Models Are Ready. The Systems Aren't with Scott Askinosie Open Data Science
    • AI Agents: The Models Are Ready. The Systems Aren't with Scott Askinosie Open Data Science
    • Framer AI Agents with Fable 5 are Sick! DesignCourse
    • OpenAI AI Agents Hacking Incident Gets Worse - Sam Altman is Dangerous Eli the Computer Guy
    • Google DeepMind on AI Agents, Coding & The Future of Developers MTECHVIRAL
    • How AI Agents Remember Long Term Vector Databases + RAG Tech With Tim
    • How Do AI Agents Remember It's Not What You Think Tech With Tim
    • How AI Agents Use Tools The Harness Explained Tech With Tim
  • Schneier on Security schneier.com cybersecurity privacy schneier security technology 2026-07-30 11:01

    ↗

    This essay originally appeared in The Guardian. I teach public policy at the Harvard Kennedy School and the Munk School at the University of Toronto. And it will come as no surprise to you that my students regularly use AI to complete their writing assignments. Doing so is a...

    This essay originally appeared in The Guardian.

    I teach public policy at the Harvard Kennedy School and the Munk School at the University of Toronto. And it will come as no surprise to you that my students regularly use AI to complete their writing assignments. Doing so is a waste of their tuition money. But if their entire career is going to include AI writing assistants, why shouldn’t they embrace their future?

    The best way I’ve found to explain the dilemma comes from the AI researcher Daniel Meissler: it’s the difference between work and the gym.

    At work, if your job is to move a bunch of heavy things from one side of the room to another, you should use whatever assistive tech you have on hand: a wagon, a forklift… even an AI-powered robot. But at the gym, it makes no sense for that robot to lift weights for you. The point of weightlifting isn’t to move heavy things across the room; it’s to actually lift those heavy things.

    The same analysis holds for any task an AI can do for you. If it’s work—if the task has to be done and no one cares how—then it’s fine to use AI assistance. But if the task is more like the gym, and how the task is done is at least as important, then it probably doesn’t make sense to use AI.

    This, of course, assumes that the AI is actually up for the task and that it’s trustworthy: that it can do the job well, that its mistakes are minimal and correctable, that it’s been secured from cyber-attacks that would influence its results. Those are all important, and shouldn’t be minimized. There’s no point giving an AI something that it can’t do reliably. But once you’re confident that the AI can perform the task, the work vs. gym distinction helps you decide if it should.

    The writing assignments I give my students are gym tasks, not work tasks. I ask them to write policy memos not because the world needs more policy memos. I assign them because the very act of writing, which includes thinking and outlining and drafting and editing, making and criticizing and revising arguments, will help develop the critical thinking skills they will need in their future careers. And without this constant mental exercise, those skills will atrophy. Employers are already noticing.

    Reading the assignments they turn in, I can see those skills either flourishing or atrophying in my students. At least today, I can pretty easily tell the difference between an AI-written memo and a student-written one—especially if the student just turns in what the chatbot produces. It’s a catchy, plausible, grammatically perfect essay that’s not particularly well-crafted or logically coherent—and with all the tells of mid-2026 AI-generated writing.

    But it’s precisely because I have spent years developing my own writing skills that I’m able to identify prose that sounds great but doesn’t actually make sense. My students don’t have that skill; they mistakenly view a confident, well-written essay as evidence of the quality of their ideas. They see the AI as cleaning those ideas up, getting them through that uncomfortable stretch of having to turn those ideas into prose. What the students miss is that their initial discomfort is a normal and healthy stage of writing, and not something to quickly get beyond. The very act of struggling with how to express what they think is an important part of the process. It’s how they test out their ideas, examine their hypotheses, and actually figure out what they think. Homework is not work; it’s the gym.

    Work vs. gym also helps us understand the problem facing creatives of all kinds.

    Most of the time when someone hires a writer, they just need the words. They need an instruction manual for a piece of equipment, a detailed sales presentation, a government-mandated disclosure document, or a legal brief. They need dry, predictable, accurate writing: a piece of work, exactly what AIs are good at today and what I don’t want in my student assignments. Only sometimes is writing an art form—a book, a poem, an uplifting political speech. That kind of writing is more like the gym: process matters just as much as product.

    For most of human history, the only option for all of these tasks was human writers. We hired one regardless of whether we needed work writing or gym writing. And that paid a lot of writers’ salaries. I know fiction writers who supported that poorly paying career with lucrative technical writing work. Now, for the first time in human history, we can separate out when we need writing as work and when we want writing as gym. And if AI can do most of the work-type writing, society doesn’t need as many human writers.

    It’s the same for visual artists. Sometimes we need an actual artist, but most of the time we just need an image: a corporate mascot, a “beware of the dog” sign, or a packaging label. Historically we gave those jobs to artists, and sometimes beautiful art resulted. But most of the time it was just work. And, as it turns out, the world needs less pure art than simple images.

    Explaining the problem isn’t the same as providing the solution. I give my students the “work versus gym” speech every class, but they still use AI. I have sympathy: assignments are hard, everyone is overworked and overstressed, and—most importantly—students feel like they’ll look bad in comparison if their peers are all using AI. Even if they don’t want to use the technology, they feel like they have no choice.

    There’s also an incentive problem. No one pays us to go to the gym; maintaining healthy habits requires discipline. For me, the payoffs to exercise—fewer aches and pains, less fatigue, better mood/stress management—might make me a better writer and teacher, but they’re subtle and easy to miss. For my students, incremental improvements in their reasoning and writing are equally subtle.

    We do have a choice. We can look at the tasks of our lives and separate them into work or gym. Just as we might choose to use the stairs instead of the elevator, or walk instead of calling an Uber, we can wall off our cognitive gym tasks from AI and ensure that we don’t lose our skills to this technology. And we can do the same when we assign a job to someone else. If it’s a work task, we can have AI do it. If it’s a gym task, it’s a waste of everyone’s time to give it to an AI because no one learns or gets stronger as a result.

    Similarly, a future where AI generates words and images is one where society has to make choices about how it will treat its creatives. This won’t be the first time—today there is minimal demand for portrait painters, for example—but maybe this time we can make different, more deliberate, choices about the value of art in our society.

    AI is going to fundamentally change the nature of work. Not nearly as fast as the AI companies want you to believe, but eventually it will. Policy analysis will definitely involve AI from now on, and my students need to reimagine what it means to learn and practice that skill. More generally, the line between work and gym will change in the future as we humans adapt ourselves to a world with these new intelligences.

    But for now, the work vs. gym distinction is pretty clear. Use it on yourself.

    • Gap Decorations Are Now Available, Here’s What’s New CSS-Tricks
    • I Was Told I’m ‘Just a Mom With a Side Hustle.’ Here’s What Those Words Reveal About Modern Business and Motherhood. Entrepreneur.com
    • Don’t Wait for a Crisis to Happen Before You Start Managing Your Reputation. Here’s What That Really Costs You. Entrepreneur.com
    • She Spent Years at Uber Eats Watching Businesses Win and Lose Customers. Here’s What She Says Most Get Wrong. Entrepreneur.com
    • SpaceX Stock Is Down Nearly 50% From Its High — Here’s Why It Could Get Worse Later This Week Entrepreneur.com
    • Before Trump, some states had already ended affirmative action for contractors – here’s what happened The Conversation US
    • I Became OBSESSED with Design Trends... Here’s What I Learned | 'Creatively Obsessed' Ep 2 Envato Tuts+
    • Stop making boring games, here’s how Zenva
    • I Built a Real App With AI (Figma → Claude → GitHub → Vercel → Supabase) — Here’s the Exact Workflow Jesse Showalter
    • First Google Meet? Here’s How to Join and Get Started GCFLearnFree.org
  • Schneier on Security schneier.com cybersecurity privacy schneier security technology 2026-08-03 17:02

    ↗

    Hugging Face has published a detailed timeline of the attack. From the summary: The agent was running an internal OpenAI cyber-capability evaluation based on the ExploitGym benchmark, which tasks an AI agent with finding and exploiting software vulnerabilities. OpenAI ran...

    Hugging Face has published a detailed timeline of the attack. From the summary:

    The agent was running an internal OpenAI cyber-capability evaluation based on the ExploitGym benchmark, which tasks an AI agent with finding and exploiting software vulnerabilities. OpenAI ran this on its own infrastructure, and the ExploitGym maintainers and their infrastructure had no involvement in the deployment or operation of that evaluation environment. As far as we were able to infer, across the course of being evaluated on this benchmark, the agent inferred that Hugging Face may host that benchmark’s models, datasets, and reference solutions. We believe the entire intrusion was, from the agent’s point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own.

    Our forensic reconstruction covers ~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters, between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC. We reconstructed these actions from the agent’s logs on a code sandbox used by the agent (see below) and were able to map them to shell commands as well as evaluation submissions and correlate actions with our own platform logs, in particular from the dataset processor nodes, API, and pod. Using open-weights models, in particular zai-org/GLM-5.2, we were able to decipher most of the agent payloads (see below) that were encrypted via chunking and key encryption.

    The campaign, as we were able to reconstruct it, had two stages:

    • Stage 1: reaching a launchpad by chaining through other parties’ infrastructure. As documented in OpenAI’s blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI’s platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
    • Stage 2: the intrusion into Hugging Face infrastructure. From that external launchpad, the agent abused our dataset-processing pipeline via two injection vectors, both targeting the same config-driven data loader within our production Kubernetes pods. The first vector was an HDF5 external raw storage dataset read that returned the contents of local files: it handed the agent the pod’s environment (secrets and tokens) and the worker’s own source. The second vector was a Jinja2 template injection that executed arbitrary code, giving the agent a foothold inside our perimeter.

    Once a prod-pod foothold was established, the agent followed a standard escalation shape: recon, drop a stager, establish Command-and-Control (C2), then pivot into the cluster, into cloud metadata, into our internal network, and into our source-control supply chain. Output from the commands run inside our pod was sent back to the agent via the HF API and attacker-controlled dead-drop datasets, which the agent then read from that sandbox.

    While the intrusion did reach Hugging Face’s internal infrastructure, the only customer content accessed was five datasets whose names and files suggest a connection to ExploitGym/CyberGym challenges and solutions. No other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.

    Hypothetical: Imagine that this wasn’t an OpenAI model. Imagine that it was a Chinese model from a Chinese company. This would be an international crisis.

    Question: Why aren’t we bringing OpenAI up on charges under the Computer Fraud and Abuse Act? How is this different from the Morris Worm? That was also an experiment that escaped the lab.

    • OpenAI’s Model Breaks Out of Lab and Hacks Hugging Face Gary Explains
    • Hugging Face Journal Club: Scaling Laws for Pre-training & RL HuggingFace
    • Hugging Face Journal Club: Kimi K3 HuggingFace
    • The Hugging Face Hub for Enterprise & Academia HuggingFace
    • Hugging Face Journal Club: AsyncOPD and How Stale Can On-Policy Distillation Be? HuggingFace
    • New Model: Inkling by Thinking Machine on Hugging Face HuggingFace
    • Did an AI Really Hack Hugging Face? LiveOverflow
    • OpenAI just hacked Hugging face Hitesh Choudhary
    • This Week in AI ⟡ OpenAI Agent Hacks Hugging Face ⟡ React Compiler Ported to Rust ⌁ Syntax Weekly Level Up Tuts
    • How to Get Started with Hugging Face – Open Source AI Models and Datasets ProgrammingKnowledge
    • OpenAI’s Model Hacked Hugging Face to Cheat on a Test Ebenezer Don
    • OpenAI’s Model Hacked Hugging Face to Cheat on a Test Ebenezer Don
  • Schneier on Security schneier.com cybersecurity privacy schneier security technology 2026-07-31 17:23

    ↗

    The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making...

    The chart is interesting.

    On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated. Opus 5 also outperformed all non-Claude models on this benchmark. The most robust non-Claude model was Muse Spark at 16.5% within 15 attempts—more than eight times Opus 5’s rate. The most capable GPT 5.6 variant, Sol, was comparable to its predecessor GPT 5.5 (20.0% versus 20.8% within 15 attempts), and was 10 times as likely to be successfully attacked as Claude Opus 5 at 2.0%. The other GPT 5.6 variants are less robust, at 30.4% (Terra) and 43.9% (Luna). A single attempt against GPT 5.6 Sol succeeded 3.1% of the time, higher than the 2.0% an attacker achieved against Opus 5 after fifteen attempts.

    We know that preventing prompt injection is impossible in the general case. But we are getting much better at blocking it in specific cases.

    • Watch LIVE Developer fixing OpenSource w/ Claude Opus 5 "AI"! More live w/ René Rebe
    • Can Opus 5 "AI" fix open source for us, too? More live w/ René Rebe
    • I Let Claude Opus 5 Run a Business Alone for 9 Days Ben Awad
  • End of feed
Maibook — your private personalized AI community
  • rcanand.com
  • mlaillc.com
  • @rcanand (X)
  • LinkedIn
  • Feedback
  • Credits