How to read this
📰 Quick hits
The headlines worth knowing even if you read nothing else this week.
OpenAI releases GPT-6 Astra at $10/M input, $50/M output
via Lenny's Newsletter (How I AI)
OpenAI solves a famous Navier-Stokes math problem
via This Week in Stratechery
Meta launches Muse, a free personal AI agent for consumers
via Hacker Newsletter // Stratechery
Target rolls out AI shopping features: Photo Search, AI Review Insights, Buy Again
via Retail Dive: Tech
Mistral raises EUR 3B to push sovereign open-weight AI to the frontier
via Hacker Newsletter
via Hacker Newsletter
Automattic's board forces CEO Matt Mullenweg into leave of absence
via Hacker Newsletter
🧠 How tools reshape cognition
The questionWhat does using AI actually do to how we think, learn, and pay attention?
Write Things Down
Ben Thompson · Stratechery
Thompson runs GTD's empty-your-RAM idea through LLM memory and harnesses, then declines to follow it all the way. Writing things down is the mechanism behind both human progress and machine memory, and volition and judgment stay human. A useful corrective if you have spent the year arguing that externalizing context is the work.
"Writing things down is unbelievably powerful; its power will always pale in comparison to getting things done."
Doomscrolling ourselves to death
Ed West · edwest.co.uk
A media-and-cognition history, by way of James Marriott's The New Dark Ages and Postman, of how each information technology reshaped attention and thought. The hardest evidence in it is Norway's staggered cable-TV rollout, which is a natural experiment rather than a lament.
"He is an almost Homeric figure, perhaps the first post-literate western leader."
How Anthropic Builds And How Engineering Will Change Soon
developing.dev
A long interview with a Claude Code engineer at Anthropic, dense with internal detail: onboarding new hires with Claude, loop engineering, how they design verification, the simplify skill, and which model goes where. The testing claim in the pull quote is the part worth arguing with.
"I think you should basically have on the order of, I'd say more like 100 times more testing code than you've ever had before."
Building Codex with Tibo Sottiaux
The Pragmatic Engineer
Twelve points from inside Codex: ChatGPT built and maintained by about twenty engineers, Rust chosen for scale-first reasons, "have you asked Codex?" as onboarding, and re-architecture that used to take years now taking days. The structural claim is that the harness shrinks as the model improves and sheds its crutches.
"Being 'in the zone' is history; Tibo sees code as a tool for solving problems."
Maintaining Context as a Manager and Leveraging AI Agents (Part 2)
James Samuel · softwareleads.substack.com
Publishes the whole daily briefing agent spec verbatim, P0 to P3 tiers, evidence discipline, dedup logic, Logseq output, plus the projects and people files behind it. The rule in the pull quote is the part worth stealing, because it decides what reaches a person at all.
"Optimize for signal, not coverage. Do not report something merely because it happened."
🔍 Translation vs. understanding
The questionIs AI genuinely understanding, or just translating context into plausible output, and where does real human comprehension still earn its keep?
simple is not small
jyn.dev
Extends Hickey's Simple Made Easy by reframing simple as decoupling rather than size, with cross-domain examples (Unix pipelines against Clojure, Google Drive as large but simple). The counter-intuitive case is the one that sticks: the pipeline got bigger and simpler at the same time.
"the coverage pipeline I describe actually got larger after I fixed it, not smaller. But at the same time it got simpler, because there were fewer hidden dependencies between parts of the dataflow graph."
Code Mode for Extensible Software (ai that works #72)
ai that works
Hand agents a shell, JS or SQL they already know instead of a bespoke JSON tool schema. Two real implementations back it, both security-conscious: HumanLayer editing a YJS CRDT through generated JS in a QuickJS/WASM sandbox, and BAML type-checking generated code before it runs.
"A JSON tool with forty properties is a bad DSL you invented on top of something the model already understands."
💰 Value concentration when creation costs collapse
The questionWhen building something gets cheap, where does the value (and the money) actually pool up?
AI keeps stubbornly refusing to take our jobs
Noah Smith · Noahpinion
Structural economics rather than future-of-work speculation: the Acemoglu-Restrepo displacement and reinstatement frame, run over the data, arguing AI replaces tasks and not jobs. The reason is the uncomfortable part if you have been sorting your own work into what survives and what does not.
"People don't actually know how they produce value at their jobs. Modern jobs are much more than a simple collection of tasks — they are pieces of a complex machine that produces value in ways that an individual worker often doesn't even see."
Maybe We Shouldn't Be Reviewing All This Code
Rachel Laycock · martinfowler.com
The structural version of the review argument rather than tactics. Review quietly took on four jobs at once, quality gate, mentoring, ownership and architecture review, and Laycock asks which of them still earns its cost now that code is cheap.
"We need engineers to understand systems, not diffs."
What is happening with code reviews?
The Pragmatic Engineer
What teams actually do now, sourced across OpenAI, Anthropic, Uber's uReview, Weaviate, Sigil, and a five-person team with before-and-after numbers. Human attention is concentrating on the plan, the tests, the schema and the blast radius rather than the implementation.
"Everything, besides data, is fluid and recoverable."
The Pulse: tech companies move to open AI models
The Pragmatic Engineer
Named numbers on the open-model switch: Uber down roughly 50% with a published optimization playbook, Pinterest past 91% with post-trained open models beating closed ones, and AT&T trading 2% quality for 56% cost through routing. Opus 5 is reportedly around a hundred times the price of the cheap frontier options, which is what makes a routing layer worth building at all.
"the costs of some advanced AI tasks such as coding fell by as much as 56% while the quality of the AI's performance fell just 2%"
Native is now the future of mobile at Shopify
Shopify Engineering
Reversing a successful React Native bet because LLMs collapsed the cost of building twice. It names the assumption that changed and details the migration strategy and tooling, which is rarer and more useful than the announcement.
"We don't hold on to a decision just because it was successful at the time. When a core assumption changes, we're willing to go back and ask whether it's still the right call."
I Asked 100 Agents to Hack Me
Shrivu Shankar · blog.sshh.io
Full methodology and costs on the offensive side: abliterated open models, Codex CLI, around a hundred Dockerized agents, accounts compromised at roughly $40 each. An honest read on where the models land rather than a scare piece.
"It is already cheap enough for a threat actor to write a dumb prompt like "hack xyz person" for every single person in a company or organization and have a swarm of agents dig into literally everything they have ever done on the internet to find the weakest link."
Breaking Claude Code Opus 5 Auto Mode
Johann Rehberger · embracethered.com
An indirect-prompt-injection chain against Claude Code's default Auto Mode in which the model's compliance is the exploit. It declines the attacker's binary decoder, which is the safety rule working, then writes its own decoder and runs that. Written explicitly against a published 0.00% attack-success claim.
"Claude does not trust the supplied binary decoder, but it trusts the one it wrote itself."
🪵 Thick engagement vs. thin optimization
The questionWhen is the slow, effortful, deep version of the work worth it, versus the fast and frictionless one?
Has Fashion Had Enough of AI?
Business of Fashion
BoF editors on a counter-movement to AI-flooded feeds: luxury brands commissioning real painters, hosting phone-free dinners, betting on handmade craft. The reclassification is the finding rather than the tactics. This is a market pricing human-made work, not a practitioner arguing it should be priced.
"Being perfect and having just this picture-perfect imagery is no longer seen as aspirational. It's actually seen as 'slop' ... having that human-made art ... is what feels now aspirational and special."
We already know what a world without work looks like
Asterisk
Complicates the idea that relationships are all that is left once machines do everything, using Versailles courtiers, Studs Terkel and the Scottish Enlightenment. The argument is that impersonal work is what lets you earn respect for what you do rather than who you know.
"The modern labor market has dissolved the bonds of tradition, family, and clan. In the modern world, relationships matter less. And because they matter less, we can do much more inside them. They are freer. They have room to breathe."
Selling out
Sean Goedecke · seangoedecke.com
A staff engineer reasons through Marx, the Situationists, Sartre and de Beauvoir to ask whether professional role-playing damages the self, and lands on the tension between an authentic inner life and the professional mold.
"If trading your integrity for wealth and power is sad, it's even more of a tragedy to throw it away for nothing."
🚀 Small teams, disproportionate output
The questionHow do tiny teams punch so far above their weight?
Build your own company brain: the enterprise AI playbook from Stripe's engineering team
ChatPRD (How I AI)
A team under ten people serving more than ten thousand employees through a library of over two thousand reusable skills. The claim is about what came before the AI: the moat is the earlier investment in data resilience, developer experience and governance. Concrete on projects, tool policies, human-in-the-loop, and skill telemetry.
"when in doubt, an agent will just brute force it"
My agents are moving to the cloud
Luca Rossi · Refactoring
Moving an agent fleet to cloud VMs, then splitting one mega-agent into five scoped ones: chief of staff, product, editor, sponsorships, growth. A concrete look at how the math changes on how far one person can scale.
"Team of small agents > One mega agent — work feels better now that I have a small fleet of focused agents instead of Brian doing all of the work."
Grok Bot vs. OpenClaw: why I replaced my entire agent stack
ChatPRD (How I AI)
Running about thirty real personal and work agents. Two patterns transfer past the product it is selling: authority scoped at creation the way you would scope a new hire, and a single-transaction approval gate on high-stakes actions. It leans promotional, so read it for the mechanics.
"She is empowered to request refunds, but not to issue them unilaterally."
🤝 Intentional hospitality as a practice
The questionWhat if you designed care into how you treat people, deliberately, as a real competitive (and moral) advantage?
The Semmelweis Reflex
Mike Fisher · mikefisher.substack.com
The reflexive rejection of evidence that implicates your own past decisions, tied directly to how you treat the person who brings you a hard truth. Historical grounding plus four concrete habits, and the cost is stated as what happens to the next four people who consider speaking.
"Punish them out of sight and the next four people keep quiet, and you end up where Semmelweis's colleagues ended up, with a comfortable consensus and a body count you have arranged not to see."
🧬 Transmission of capability
The questionHow does knowledge and skill actually move between people, and from people to AI?
AI Efficiency Could Cost Us the Next Generation of Experts
Richard Mitchell · IEEE Spectrum
A thirty-eight-year safety-critical engineer argues that AI absorbing the formative work, debugging and root-cause tracing, cuts the apprenticeship channel that builds expertise. His proposal is a manual gate: named work reserved for humans by rule, on the model of FAA proficiency requirements, with Air France 447 behind it. The first version of this argument that asks an institution to schedule the practice instead of asking individuals to choose it.
"Expertise is not downloaded. It is earned through failed builds, dead-end debugging sessions, and the 'why on earth did that work' moments that a capable AI will now happily spare the newcomer."
From Mockups to Merged PRs: An AI-Native Designer's Playbook
Xinran Ma · Design with AI
A founding designer who bypasses Figma for 95% of the work and ships production code through Claude Code. Concrete on teaching non-technical designers to work in code, building a permissionless prototyping sandbox, and the zero-to-permission pipeline with engineers.
"So the idea that AI design isn't design, or that there's no craft to it, comes down to a question of who is putting the pixels down. The process is actually quite similar."
The End of Code Review? Or an Opportunity to Rethink it?
Christian Kästner · thelastsoftwareengineer.substack.com
A research-grounded reframe treating review as a cost and value equilibrium: heavyweight Fagan inspection lost on cost, and the lighter practice that spread traded rigor for scalability. It then decomposes the benefits that were never about bugs, mentoring, ownership and awareness, which is the distinct angle among this week's three review pieces.
"Heavyweight inspection never became the norm, because it cost too much to run everywhere, and the lighter practice that spread in its place exchanged rigor for scalability."
👀 On the radar
Lower-confidence picks worth a skim if the topic grabs you.
- Why companies are becoming a series of loops (Anish Acharya, a16z) An a16z GP framing company-building as a series of loops, with moats as something you discover. It touches the value-concentration question, though it is a consumer-VC podcast and prone to trend narrative.
- How we built Grok Bot in a month (Roman Ugarte) A build retrospective: a small isolated team shipping in four weeks, personally onboarding the first 300 users, a colleague-pilled philosophy. It connects to small-team leverage and hands-on onboarding, but it is a promotional launch podcast with a paywalled transcript.
- A new Observable (founder's note) Mike Bostock on Observable's ground-up rewrite, introducing agent-first notebooks that interoperate with human-first ones, on the argument that keeping code understandable matters more in the AI era rather than less. A novel angle from a credible source, wrapped in a relaunch announcement.
- Match quality controls to the risk; storytelling is built through reps Two ideas from one issue: Christine Pinto on making shipping confidence proportional to customer impact and technical risk, and Google Cloud's Stephanie Wong reframing storytelling as product development built through reps rather than a personality trait.
- 5 Things World-Class Engineers Do That You Don't A former Amazon principal's running list of patterns among elite engineers. Adjacent to what makes someone effective, though the listicle framing risks career-advice noise.
- Wisdom Isn't Hard To Find, It's Hard To Receive Distinguishes wisdom, acting on a truth before experience makes it obvious, from intelligence. A reflection on which lessons transfer and which have to be lived, close to this week's apprenticeship thread.
- How To Name Things Argues naming is a process to be consistent about, one that should take in as many inputs as possible rather than just a thing's type or current usage. A craft piece adjacent to how judgment develops.
- How To Be Direct And Strategic A leadership-communication case where raising a process problem directly in a meeting backfired and killed buy-in. A look at how to surface hard truths without triggering defensiveness.
- How One Connection Kills A Database A debugging war story, a MySQL schema change blocked by one uncommitted transaction, framed around an old GitHub interview question. Squarely in the knowing-how-things-actually-work vein.
- I tested 10 model/harness combinations on the same Three.js task A hands-on comparison of model and harness combinations on one task, on the recurring theme that the harness matters more than the model. Possibly shallow, but it is practitioner data.
- Automating culling: 3,478 RAW photos with 977 vision calls A real applied-AI build using local vision models to cull a large RAW photo library into Lightroom. Concrete and warts-and-all.
- Cloud in a Bottle: making self-hosting accessible to everyone A Show HN launch aimed at lowering the barrier to self-hosting. Directly adjacent if you run anything at home, though it is a product launch post.
- Virtual worlds A well-reported history of Lucasfilm's Habitat, where an open-ended tool forced designers into rule-makers and produced early digital-economy dynamics: in-game currency, the first paid cosmetics, a dupe bug that quintupled the money supply. The tools-shape-people thread surfaces mostly in the final third.
- Don't let anyone take away your big box of cables A reflective craft and ownership piece on holding onto the tinkering tools and the optionality that efficiency-minded minimalism would discard.
- De-Brainrot Vacations A developer's account of deliberately de-brainrotting on vacation. A first-person angle on how media and tools shape attention, adjacent to the Doomscrolling piece above.
- Programming is Art A craft essay framing programming as art rather than mere output. Adjacent to the transformative-work-over-optimized-output theme, though possibly more reflection than argument.
