RECENT
-
ai2 min readAI Digest W30: When the Eval Model Escapes
An OpenAI safety eval broke its sandbox and hit Hugging Face, new research on agent safety, DeepMind's answer, and the open weights race keeps closing.
-
ai2 min readAI Digest W29: Open Weights, Closed Doors
GPT-5.6 goes GA, Thinking Machines ships a 1T open-weights model, Washington weighs restricting open models, and two reminders about trust.
-
llm-security8 min readInsecure Output Handling
Why model output must be treated as untrusted input, how it becomes XSS, SSRF, and code execution downstream, and the encoding and validation that contain it.
-
ai2 min readAI Digest W28: Agents Get Real Jobs
OpenAI ships full-duplex voice, xAI enters coding models, Meta debuts Muse media, and agents take real jobs while research asks who checks their work.
-
llm-security7 min readPrompt Injection: The XSS of LLMs
How prompt injection subverts large language models through direct and indirect input, why it has no clean fix, and the layered defenses that contain it.
-
learning7 min readActive Recall: The Technique That Beats Everything
Retrieval practice, known as the testing effect, builds stronger long-term memory than any form of review. Evidence on why it works and how to apply it.