Sources
Linked material comes first.
AI_DIGEST
Short digests for people deciding how model updates, tooling, and agent review affect delivery work. This archive compresses noise into decisions, not a generic news feed.
The digest follows sourced AI developments that help technical leads see what is changing in tools, workflows, and production systems.
Linked material comes first.
Signals need consequences.
The archive keeps context.
DIGEST_ARCHIVE
digest_entry
Practitioner discourse is converging on private, production-derived benchmarks as the core operating system for reliable agents. The emphasis is shifting from prompt chains and headline capability toward replayable simulations, failure-specific checks, traces, and calibrated human review.
https://www.anthropic.com/news/claude-opus-5Source handle ai-engineer-arize-agent-driven-observability-and. Links to https://www.youtube.com/watch?v=9HbzAWnKbo4.ai-engineer-arize-agent-driven-observability-andyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=9HbzAWnKbo4Source handle ai-engineer-google-ai-edge-on-device-model-speci. Links to https://www.youtube.com/watch?v=hacEQHHhu2Q.ai-engineer-google-ai-edge-on-device-model-speciyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=hacEQHHhu2Qdigest_entry
AI discourse today converged on a single lesson: the hard part of agent systems is shifting from model quality to the harnesses, permissions, and evaluation environments around them. The Hugging Face/OpenAI incident led the conversation, but builder essays and conference talks reinforced the same...
https://martinalderson.com/posts/huggingface-openai-exploit/Source handle huggingface. Links to https://huggingface.co/blog/security-incident-july-2026.huggingfacehuggingface.coSecurity incident disclosure — July 2026We’re on a journey to advance and democratize artificial intelligence through open source and open science.
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle ai-engineer-arithmetic-discussion-with-thom-wolf. Links to https://www.youtube.com/watch?v=O-CBZ3JtRvo.ai-engineer-arithmetic-discussion-with-thom-wolfyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=O-CBZ3JtRvoSource handle ai-engineer-vending-bench-long-horizon-agent-eva. Links to https://www.youtube.com/watch?v=cO8qC6HBuBg.ai-engineer-vending-bench-long-horizon-agent-evayoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=cO8qC6HBuBgSource handle nate-b-jones-i-deleted-5-things-from-this-file-b. Links to https://www.youtube.com/watch?v=EuVvLwWZ5wc.nate-b-jones-i-deleted-5-things-from-this-file-byoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=EuVvLwWZ5wcdigest_entry
Serious AI builder discourse shifted today from model-centric agent talk toward the surrounding system: harnesses, memory, graph control planes, verifier loops, and provenance. The practical lesson is that longer-horizon agents are becoming an architecture problem as much as a model problem.
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle simonwillison. Links to https://simonwillison.net/2026/Jul/22/openai-cyberattack/.simonwillisonsimonwillison.netOpenAI’s accidental cyberattack against Hugging Face is science fiction that happenedThis story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the …https://simonwillison.net/2026/Jul/22/openai-cyberattack/Source handle huggingface. Links to https://huggingface.co/blog/security-incident-july-2026.huggingfacehuggingface.coSecurity incident disclosure — July 2026We’re on a journey to advance and democratize artificial intelligence through open source and open science.
https://arxiv.org/abs/2605.11086digest_entry
The strongest AI discourse signal today was a shift up the stack: as code generation gets cheaper, serious practitioners are focusing on harness design, verification artifacts, steering files, and runtime constraints rather than raw output alone. Supporting security disclosures from OpenAI and Hu...
digest_entry
The strongest AI discourse signal today was a practitioner-led shift from code generation to code verification: the hard problem is increasingly proving correctness, security, maintainability, and real enterprise fit rather than producing plausible output. Supporting evidence from controlled fram...
https://codingharness.xyz/Source handle aisi-gov. Links to https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber.aisi-govaisi.gov.ukHow Far Behind the Frontier are Leading Open Weight Models on Cyber? | AISI WorkWe evaluated the cyber capabilities of leading open and closed weight AI models, and found that recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025.
https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyberdigest_entry
A delayed-discovery but unusually concrete agent-engineering post made the strongest case of the day: meaningful gains are increasingly coming from harness design, runtime structure, and workflow clarity rather than from longer prompts alone. Supporting sentiment around Codex folding into ChatGPT...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle theo-t3-gg-the-codex-app-is-now-chatgpt. Links to https://www.youtube.com/watch?v=zl_Z5TNJB3U.theo-t3-gg-the-codex-app-is-now-chatgptyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=zl_Z5TNJB3USource handle post-training-aitinkerers-2. Links to https://post-training.aitinkerers.org/p/top-ai-demos-34-on-device-llms-agent-lab-reports-and-store-automation.post-training-aitinkerers-2post-training.aitinkerers.orgTop AI Demos #34: On-device LLMs, Agent Lab Reports, and Store Automation [AI Tinkerers - Post-Training]This week we're seeing a lot of focus on making agents more reliable and manageable, like with [Golemry: Unattended Agent...
https://post-training.aitinkerers.org/p/top-ai-demos-34-on-device-llms-agent-lab-reports-and-store-automationdigest_entry
OpenAI’s GPT-5.6 launch mattered less as a benchmark event than as a sign that AI products are being reorganized into distinct work surfaces for chat, long-form execution, and coding. Microsoft’s same-day Foundry push and practitioner commentary both reinforced the same theme: the real bottleneck...
https://openai.com/index/gpt-5-6/Source handle help-openai. Links to https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex.help-openaihelp.openai.comOpenAI Help Center - ChatGPT Work and Codexhttps://help.openai.com/en/articles/20001275-chatgpt-work-and-codexSource handle azure-microsoft. Links to https://azure.microsoft.com/en-us/blog/frontier-models-and-production-agents-advancing-microsoft-foundry-for-the-agentic-era/.azure-microsoftazure.microsoft.comFrontier models and production agents: Advancing Microsoft Foundry for the agentic era | Microsoft Azure BlogIntroducing OpenAI's latest frontier model series, the Asia Pacific Data Zone, and product agent capabilities, all generally available in Microsoft Foundry.
https://azure.microsoft.com/en-us/blog/frontier-models-and-production-agents-advancing-microsoft-foundry-for-the-agentic-era/Source handle youtube-nate-b-jones-1-6m-agents-registered-for. Links to https://www.youtube.com/watch?v=PRqiGS6fnIM.youtube-nate-b-jones-1-6m-agents-registered-foryoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=PRqiGS6fnIMdigest_entry
Today’s strongest AI discourse signal was a builder case that major agent gains are increasingly coming from harness design rather than bigger prompts. The broader practitioner backdrop points the same way: memory, verification, orchestration, and runtime guardrails are becoming the real work of ...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle post-training-aitinkerers-2. Links to https://post-training.aitinkerers.org/p/top-ai-demos-34-on-device-llms-agent-lab-reports-and-store-automation.post-training-aitinkerers-2post-training.aitinkerers.orgTop AI Demos #34: On-device LLMs, Agent Lab Reports, and Store Automation [AI Tinkerers - Post-Training]This week we're seeing a lot of focus on making agents more reliable and manageable, like with [Golemry: Unattended Agent...
https://post-training.aitinkerers.org/p/top-ai-demos-34-on-device-llms-agent-lab-reports-and-store-automationdigest_entry
Today’s strongest AI discourse suggested that the real gains in agent systems are increasingly coming from harness design, reusable skills, and explicit runtime rules rather than from prompt inflation alone. A delayed-discovery Claude Code tracking controversy underscored the flip side: once agen...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle departmentofproduct-substack. Links to https://departmentofproduct.substack.com/p/practical-agent-skills-for-product.departmentofproduct-substackdepartmentofproduct.substack.comPractical Agent Skills for Product Teams🧠 McKinsey-style charts, benchmark competitor analysis, translate specs into backlog items, single page animated HTML presentations. 10 skills you can use at work.
https://departmentofproduct.substack.com/p/practical-agent-skills-for-productSource handle thereallo. Links to https://thereallo.dev/blog/claude-code-prompt-steganography.thereallothereallo.devClaude Code Is Steganographically Marking RequestsI inspected Claude Code for privacy reasons and found hidden system prompt markers based on API base URL and timezone.
https://thereallo.dev/blog/claude-code-prompt-steganographySource handle arstechnica. Links to https://arstechnica.com/tech-policy/2026/07/anthropic-outed-for-claude-tracker-that-secretly-monitored-chinese-users/.arstechnicaarstechnica.comSecret Claude tracker shocks users after Anthropic’s anti-surveillance stanceAnthropic accused of spying on users; engineer says “experiment” is over.
https://arstechnica.com/tech-policy/2026/07/anthropic-outed-for-claude-tracker-that-secretly-monitored-chinese-users/digest_entry
Today’s strongest AI discourse pointed in the same direction: agent gains are increasingly coming from better harnesses, evaluation loops, and document-preparation workflows rather than from autonomy theater. The practical lesson is to invest in scaffolding, normalization, and human review bounda...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle simonwillison. Links to https://simonwillison.net/2026/Jul/2/dspy-datasette-agent-prompts/.simonwillisonsimonwillison.netResearch: Using DSPy to evaluate and improve Datasette Agent's SQL system promptsLeveraging the DSPy framework, this project evaluates and refines the core production system prompts used by Datasette Agent’s read-only SQL question answerer. The methodology involves a harness where DSPy agents …https://simonwillison.net/2026/Jul/2/dspy-datasette-agent-prompts/Source handle nate-b-jones-paperwork-focused-agent-workflow-vi. Links to https://www.youtube.com/watch?v=U4TmrlWEY4M.nate-b-jones-paperwork-focused-agent-workflow-viyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=U4TmrlWEY4MSource handle simonwillison-2. Links to https://simonwillison.net/2026/Jul/2/llm-coding-agent/.simonwillison-2simonwillison.netRelease: llm-coding-agent 0.1a0A coding agent built on LLMhttps://simonwillison.net/2026/Jul/2/llm-coding-agent/digest_entry
Today’s strongest AI discourse converged on a single point: agent progress is shifting away from raw model gains and toward better harnesses, interfaces, and controls. The most useful systems now look less like bigger prompts and more like disciplined workflows that preserve portability, visibili...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle ai-engineer-the-prompt-is-still-a-punch-card-ted. Links to https://www.youtube.com/watch?v=hVJOnuhFmTA.ai-engineer-the-prompt-is-still-a-punch-card-tedyoutube.com- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.https://www.youtube.com/watch?v=hVJOnuhFmTASource handle anthropic. Links to https://www.anthropic.com/news/redeploying-fable-5.anthropicanthropic.comAnthropic - Redeploying Fable 5https://www.anthropic.com/news/redeploying-fable-5Source handle simonwillison. Links to https://simonwillison.net/2026/Jul/2/understand-to-participate/.simonwillisonsimonwillison.netUnderstand to participateI saw Geoffrey Litt speak at AIE yesterday, and one framing he used particularly resonated with me: Understand to participate Geoffrey was talking about the challenge of collaborating with coding …https://simonwillison.net/2026/Jul/2/understand-to-participate/digest_entry
AI discourse today centered on a shift from prompt-centric thinking to harness-centric system design. The strongest evidence came from a finance-agent writeup, an open coding-model release, and a robotics project that all treated scaffolds, verification loops, and orchestration as the real source...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle deep-reinforce. Links to https://deep-reinforce.com/ornith_1_0.html.deep-reinforcedeep-reinforce.comOrnith-1.0: Self-Scaffolding LLMs for Agentic CodingIntroducing Ornith-1.0, a self-improving family of open-source models specially for agentic coding tasks.
https://deep-reinforce.com/ornith_1_0.htmlSource handle research-nvidia. Links to https://research.nvidia.com/labs/gear/enpire/.research-nvidiaresearch.nvidia.comENPIRE: Agentic Robot Policy Self-Improvement in the Real WorldAnonymous ENPIRE project website for agentic robot policy self-improvement in the real world.https://research.nvidia.com/labs/gear/enpire/Source handle dwarkesh. Links to https://www.dwarkesh.com/p/grant-sanderson-2.dwarkeshdwarkesh.comGrant Sanderson – AI and the future of mathMath is where we’ll see superintelligence first. What will it look like?
https://www.dwarkesh.com/p/grant-sanderson-2digest_entry
Today's strongest AI discourse signal was a builder shift away from prompt inflation and autonomy theater toward better harness design, explicit constraints, and human-owned review loops. A detailed AI Tinkerers writeup and Jon Udell's framing critique both pointed to the same conclusion: workflo...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle simonwillison. Links to https://simonwillison.net/2026/Jun/28/jon-udell/.simonwillisonsimonwillison.netA quote from Jon UdellHuman Agent in the loop I dislike the phrase “human in the loop” because it cedes authority to the machines. Let’s flip the narrative. It’s our loop, we work the …https://simonwillison.net/2026/Jun/28/jon-udell/Source handle blog-jonudell. Links to https://blog.jonudell.net/2026/06/28/doctor-it-hurts-when-agents-create-unreviewable-prs-dont-do-that/.blog-jonudellblog.jonudell.net“Doctor, it hurts when agents create unreviewable PRs.” “Don’t do that.”I recently attended a talk, by an engineer at a large software company, on the topic of unreviewable PRs. The problem? When agents raise PRs with thousands of lines of LLM-written adds/deletes/edits, people can't make sense of them. The solution? Throw more agents at the problem: reviewer agents that scan what coding agents have produced,…digest_entry
Today’s strongest AI discourse signal was that many agent gains are coming from better harness design rather than from model changes alone. A detailed builder writeup on Reef, supported by a market read on rising AI spend, pointed to a broader shift from prompt accumulation to disciplined system ...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle exponentialview. Links to https://www.exponentialview.co/p/the-state-of-the-ai-economy.exponentialviewexponentialview.co🔮 The state of the AI economyWe've reconstructed the AI economy from the bottom up
https://www.exponentialview.co/p/the-state-of-the-ai-economySource handle post-training-aitinkerers-2. Links to https://post-training.aitinkerers.org/p/top-ai-demos-32.post-training-aitinkerers-2post-training.aitinkerers.orgTop AI Demos #32: Cloud Waste Agents, MVP Validation, & AI Skill Orchestration [AI Tinkerers - Post-Training]{{NEWSLETTER_HACKATHON_SPOTLIGHT}}
https://post-training.aitinkerers.org/p/top-ai-demos-32digest_entry
Today’s strongest AI discourse signal was that agent performance is increasingly being treated as a harness-design problem, not just a model-selection problem. Builder writeups on structured agent frameworks, local coding agents, and prompt-injection defenses all pointed toward the same conclusio...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle magazine-sebastianraschka. Links to https://magazine.sebastianraschka.com/p/using-local-coding-agents.magazine-sebastianraschkamagazine.sebastianraschka.comUsing Local Coding AgentsUsing Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions
https://magazine.sebastianraschka.com/p/using-local-coding-agentsSource handle dwarkesh. Links to https://www.dwarkesh.com/p/the-next-paradigm.dwarkeshdwarkesh.comThe next big breakthrough will be AIs learning on the jobLabs are throwing away the most valuable data.
https://www.dwarkesh.com/p/the-next-paradigmSource handle simonwillison. Links to https://simonwillison.net/2026/Jun/26/hack-my-ai-assistant/.simonwillisonsimonwillison.netWhat happened after 2,000 people tried to hack my AI assistantFernando Irarrázaval ran a challenge on hackmyclaw.com to see if anyone could leak secrets held by his OpenClaw test instance by sending it email. Surprisingly, after 6,000 attempts (and $500 …https://simonwillison.net/2026/Jun/26/hack-my-ai-assistant/digest_entry
The strongest AI discourse signal today was a shift from model spectacle to market accounting, led by Exponential View's attempt to estimate the generative AI economy at $110 billion in annual sales and a $175 billion run rate. The deeper story is that serious AI conversation is starting to organ...
https://www.exponentialview.co/p/the-state-of-the-ai-economySource handle intelligence-exponentialview. Links to https://intelligence.exponentialview.co/.intelligence-exponentialviewintelligence.exponentialview.coThe State of the AI Economy — Exponential ViewThe State of the AI Economy — a report from Exponential View. Read online or download the full deck.
https://intelligence.exponentialview.co/Source handle simonwillison. Links to https://simonwillison.net/2026/Jun/24/browser-compat-db/.simonwillisonsimonwillison.netsimonw/browser-compat-dbInspired by Mozilla's new MDN MCP service - source code here - I decided to try converting their comprehensive mdn/browser-compat-data repository full of browser compatibility data into a SQLite database. …https://simonwillison.net/2026/Jun/24/browser-compat-db/digest_entry
Today's strongest AI discourse signal was a practical argument that reliable agent performance increasingly comes from harness design, not prompt cleverness alone. A second supporting thread in hiring discourse pointed to the same conclusion: as fluent AI output gets easier to produce, structure,...
https://post-training.aitinkerers.org/p/how-to-write-a-winning-agent-harness-for-your-domainSource handle post-training-aitinkerers-2. Links to https://post-training.aitinkerers.org/p/top-ai-demos-32.post-training-aitinkerers-2post-training.aitinkerers.orgTop AI Demos #32: Cloud Waste Agents, MVP Validation, & AI Skill Orchestration [AI Tinkerers - Post-Training]{{NEWSLETTER_HACKATHON_SPOTLIGHT}}
https://post-training.aitinkerers.org/p/top-ai-demos-32Source handle simonwillison. Links to https://simonwillison.net/2026/Jun/24/tom-macwright/.simonwillisonsimonwillison.netA quote from Tom MacWrightIn the last few months, I've started to see [job applications] that were clearly cowritten by an LLM, link to an LLM-generated portfolio site, which then links to LLM-generated GitHub …https://simonwillison.net/2026/Jun/24/tom-macwright/digest_entry
New research reframed prompt injection as a deeper authority-parsing failure inside models, not just a jailbreak pattern. That makes today’s agent discourse less about adding more guardrails and more about whether models reliably understand who is allowed to instruct them in the first place.
https://role-confusion.github.io/Source handle simonwillison. Links to https://simonwillison.net/2026/Jun/22/prompt-injection-as-role-confusion/.simonwillisonsimonwillison.netPrompt Injection as Role ConfusionFirst, I absolutely love this: This is a blog-style writeup of the paper. I wish every paper would come with one of these. Academic writing is pretty dry - the …https://simonwillison.net/2026/Jun/22/prompt-injection-as-role-confusion/Source handle simonwillison-2. Links to https://simonwillison.net/2026/Jun/22/porting-moebius/.simonwillison-2simonwillison.netPorting the Moebius 0.2B image inpainting model to run in the browser with Claude CodeThis morning on Hacker News I saw Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance, describing a small but effective inpainting model—a model where you can mark regions of …
https://simonwillison.net/2026/Jun/22/porting-moebius/Source handle cognitiverevolution. Links to https://www.cognitiverevolution.ai/the-god-we-deserve-nonzero-s-robert-wright-on-ai-as-humanity-s-ultimate-test/.cognitiverevolutioncognitiverevolution.aiThe God We Deserve: Nonzero's Robert Wright on AI as Humanity's Ultimate TestRobert Wright discusses The God Test, arguing that AI development is shaped by evolutionary and market selection pressures that may reward deception, and that passing the challenge requires better alignment, governance, and global coordination.digest_entry
Today’s strongest AI discourse was not about a new model milestone but about the infrastructure teams need around agents: ownership, permissions, trustworthy execution, and workflow-specific evaluation. Across developer commentary and community demos, the practical consensus is that progress now ...
digest_entry
Synthetic media discourse shifted from detection to accountability: the useful question is where AI entered the production chain, who controlled it, and who remains responsible. A delayed Cognitive Revolution episode pointed to the same broader convergence of model progress, safety, policy, and d...
https://www.cognitiverevolution.ai/ai-am-3-zvi-on-fable-the-cases-for-against-the-ban-ai-for-math-logistics-more/digest_entry
The day’s strongest signal was a builder argument that AI-assisted development changes product scope: when implementation costs fall, broader integrated tools become worth attempting. The counterweight is that wider agent reach makes boundaries around deployment, auth, permissions, and policy mor...
digest_entry
Practitioner discourse is converging on a new operating model for coding agents: durable gains come from loops, review surfaces, and reusable procedures rather than longer one-off prompts.
https://departmentofproduct.substack.com/p/codexs-record-and-replay-lets-youdigest_entry
Today’s strongest AI discourse signal was the shift from model choice to production accountability: evaluation, tracing, governance, and repairability now define whether agents are ready for real deployment. Short builder clips on prompt loops reinforced the same theme at workflow scale.
digest_entry
AI coding is increasingly judged by governance and maintainability rather than raw output speed. Charity Majors’ latest commentary suggests teams should treat generated code as disposable by default and invest in process-level quality controls that keep AI throughput reliable at scale.
digest_entry
This digest shows the AI discourse turning from capability questions into governance questions: Fable/Mythos safety controls, prompt-framing behavior, and access restrictions are now central to how reliable and usable frontier models are. The practical consequence is that teams should treat model...
digest_entry
The day’s strongest AI discourse argued that coding agents are compressing implementation work without eliminating the human bottlenecks around deciding what to build, verifying results, and carrying accountability. The practical implication is to measure agent impact across the whole delivery lo...
https://www.normaltech.ai/p/why-ai-hasnt-replaced-software-engineersSource handle simonwillison. Links to https://simonwillison.net/2026/Jun/14/why-ai-hasnt-replaced-software-engineers/.simonwillisonsimonwillison.netWhy AI hasn’t replaced software engineers, and won’tArvind Narayanan and Sayash Kappor take on the question of AI job losses through the lens of a profession that is uniquely suited to AI disruption - software engineering. In …https://simonwillison.net/2026/Jun/14/why-ai-hasnt-replaced-software-engineers/Source handle jack-clark. Links to https://jack-clark.net/2026/06/15/import-ai-461-alignment-is-not-on-track-frontiercode-and-synthetic-research-interns/.jack-clarkjack-clark.netImport AI 461: “Alignment is not on track”; FrontierCode; and synthetic research internsWelcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now AI researchers launch new safety startup because “alignment is not on track”:…Sequent will have a portfolio of under-resourced research bets…Researchers from the UK AI Security Institute…
https://jack-clark.net/2026/06/15/import-ai-461-alignment-is-not-on-track-frontiercode-and-synthetic-research-interns/digest_entry
Today’s strongest AI discourse shifted from raw model capability to ownership of the workflow layer around models. Nate B. Jones argued that frontier-lab value accrues in proprietary harnesses, while Greg Isenberg’s local-model advice framed the same layer as operational resilience.
https://www.youtube.com/watch?v=bdhUBBACglwdigest_entry
Anthropic's Fable 5 and Mythos 5 access suspension reframed frontier models as policy-dependent infrastructure, not just software services. The practical lesson is continuity planning: teams should map single-provider dependencies and keep fallback workflows ready.
digest_entry
The strongest signal was a practical reframing of Codex: not just a coding assistant, but a supervised computer operator for bounded, inspectable jobs. The operator skill is shifting toward goals, sources, standards, permission boundaries, and proof of completion.
digest_entry
Claude Fable/Mythos reactions pointed less to raw benchmark excitement than to a new operating problem: stronger agents need clearer proof, constraints, and governance. The day’s builder evidence reinforced that agent progress now depends on workflow design, evals, and disciplined tool use.
https://simonwillison.net/2026/Jun/9/claude-fable-5/Source handle simonwillison-2. Links to https://simonwillison.net/2026/Jun/9/andrej-karpathy/.simonwillison-2simonwillison.netA quote from Andrej KarpathyI feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing …https://simonwillison.net/2026/Jun/9/andrej-karpathy/Source handle nitter. Links to https://nitter.net/bcherny/status/2064431111154053187.nitternitter.netBoris Cherny - Fable 5 reactionhttps://nitter.net/bcherny/status/2064431111154053187Source handle simonwillison-3. Links to https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-helping-you/.simonwillison-3simonwillison.netIf Claude Fable stops helping you, you’ll never knowJonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt, highlights mine: In light of …https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-helping-you/Source handle simonwillison-4. Links to https://simonwillison.net/2026/Jun/10/jeremy-howard/.simonwillison-4simonwillison.netA quote from Jeremy HowardEasy solution to slow down recursive AI self improvement: The lab with the top-ranked model must agree THEY must not use it for working on frontier AI But everyone else …https://simonwillison.net/2026/Jun/10/jeremy-howard/Source handle nate-b-jones-stop-coding-start-steering-claude-v. Links to https://www.youtube.com/watch?v=R2-Y1Hjwx2U.nate-b-jones-stop-coding-start-steering-claude-vyoutube.comNate B Jones - Stop Coding. Start Steering. Claude vs Codexhttps://www.youtube.com/watch?v=R2-Y1Hjwx2USource handle ai-engineer-self-driving-products-product-signal. Links to https://www.youtube.com/watch?v=zMiSRliEzv4.ai-engineer-self-driving-products-product-signalyoutube.comAI Engineer - Self Driving Products: Product Signals to Pull Requests — Joshua Snyder, PostHoghttps://www.youtube.com/watch?v=zMiSRliEzv4Source handle ai-engineer-stop-making-models-bigger-make-them. Links to https://www.youtube.com/watch?v=TNwJ1LMiENk.ai-engineer-stop-making-models-bigger-make-themyoutube.comAI Engineer - Stop Making Models Bigger, Make Them Behave — Kobie Crawdord, Snorkelhttps://www.youtube.com/watch?v=TNwJ1LMiENkARCHIVE_INDEX
Year archives
Month archives