Agents Installed Code Nobody Owns
Researchers found 227 install commands across 120 corporate domains pointing at packages that did not exist, registered a few, and had phone-homes from Fortune 500 networks within hours. The agents did what the vendor docs told them to do, and so did the humans supervising them.
Top-Line Summary
Agents are reading vendor documentation as ground truth and installing whatever it tells them to install. Researchers found 227 install commands across 120 corporate domains pointing at packages nobody owned, registered a few of them, and started getting beacons from Fortune 500 networks within hours. Elsewhere: OpenAI cuts Cursor off in November over terms-of-service violations after SpaceX bought it, AWS commits to another 2 million NVIDIA GPUs, and GLM 5.3 Flash lands 10 points below frontier on the intelligence index at roughly nine cents a task instead of four dollars.
Show Video
The “Pre-Show” Context
A storm sat on top of the recording. Brett’s phone was going off every 15 minutes with alerts, it was dark enough at five o’clock to look like the middle of the night, and thunder cracked right as they started, which he described as high-risk podcasting. He also admits he spent the afternoon listening to the show’s intro song on a loop while prepping notes and came very close to letting the whole thing play through before they said a word.
The Engineering Rundown
-
AWS and NVIDIA to deliver 2 million additional GPUs (01:46)
Two million more NVIDIA GPUs across AWS global infrastructure starting in 2027, on top of the million AWS already committed to back in 2026, which puts the number at three million. Vera CPU-based infrastructure appears to be going fleet-wide rather than landing in one instance family, and Trainium picks up NVLink Fusion and custom HBM. The detail that stopped both of them was 100,000 of those GPUs going to US government AI factories at IL6 and above, which Brett filed under “sounds cool, slightly scary, I don’t think I like it.” The practical takeaway is smaller and more depressing: this is why nobody is buying a GPU again, and why everything costs what it costs.
-
Claude Fable 5.1 is now available on AWS (04:06)
Fable 5.1 is a covered model, which in practice means your prompts and the model’s completions are retained for at least 30 days because the tasks it is good at are the tasks Anthropic would rather you did not do. Eligible AWS customers can keep that retained data inside AWS rather than at Anthropic, controlled by a mode set on the account or on the Bedrock project. What “eligible” means is not spelled out anywhere Brett could find, and when he went to poke at Fable on Bedrock he got errors and a big red nope, lost interest, and wandered off. Travers has used 5.1 outside Bedrock and rates it noticeably better on research and data science work than anything else he has thrown at those problems.
-
Manage agents, tools and skills at scale with AWS Agent Registry (06:42)
A central catalog for agents, MCP servers and skills with ownership, versioning, an approval workflow and everything logged to CloudTrail because it is all API calls. Consumption pricing with a free tier, live in five regions. The part worth having is the auto-discovery of agents already running across the organization, because inside a large enterprise four people have already built the agent you are about to build and nobody knows. Brett’s summary of the week’s AWS news: they buy the compute, rent you the models, and sell you a registry to keep track of what you built.
-
Cognito machine-to-machine authorization (08:15)
Travers flagged this one. Yet another Cognito auth feature, but a useful one: you can now establish machine-to-machine credentials without needing a user pool, so your service accounts stop being fake humans sitting alongside real ones. This lands directly on a problem the show comes back to constantly, which is how you get an agent the credentials it needs for an account without handing over your own. Brett’s current answer is 1Password; other people use the OS keychain. Use the get-client-token API for this one.
-
SnapStart for container image functions in Lambda (09:41)
Lambda container image functions now get SnapStart, with sub-second startup. If you are packaging Lambda functions as containers, this is a cold-start improvement you get for free by turning it on. Short segment, and neither of them has run it in anger yet.
-
Personal projects: a Reddit for your second brain, and a pinball port (10:41)
Travers got tired of clicking through articles in his notes to prep for the show, so he had an agent build him a browser that renders his second brain as a Reddit-style feed, and used it to prep this episode. It took about 20 minutes. He is also porting a pinball rendering and simulation library from Unity to Godot, which is at the stage where the ball rolls and the paddles hit it. Brett is still dumping links into Obsidian and clicking them one at a time like an animal, and says he needs something cooler.
-
The Claude Code saga continues (12:40)
This is the running story of the last few weeks and it got worse. Skills stopped firing in auto mode; switching to manual brought them back but demanded approval for
ls; switching back to auto kept them working for a couple of days and then his AWS CLI commands started getting denied by the classifier. The tool told him to add the exact commands back to his settings file, and when he pushed back it produced the classic “you’re right, my explanation was wrong” after he pointed out that yesterday’s session made 363 bash calls with one classifier denial and all 19 AWS calls passed. This is hitting real work: he has a locked-down docs-only IAM role driving scheduled GitLab jobs that generate as-built documentation, it had been working fine, and it just stopped. He downloaded Codex again, finished the day in Kiro, and is talking about moving to Pi. His read is that it is the harness rather than the model, and since Claude Code is not open source there is no way to confirm that. Travers thinks the models are being RL’d for more capability and the classifier is compensating, which is a reasonable theory that does not explain why skills get ignored. -
NVIDIA in talks to acquire Hugging Face for $13 billion (23:09)
Reported, not signed, and it would be NVIDIA’s largest acquisition. Neither of them can work out the why. Stripe buying OpenRouter made sense because you take a slice of every token that runs through it and you are already the payment platform; buying the GitHub of models is harder to justify at that price. The best theory they land on is control over which architectures the main distribution point for models and datasets supports, which matters more now that GLM 5.3 Flash was trained and served entirely on Chinese hardware. Wild speculation, and they say so.
-
OpenAI’s decision on Cursor following its acquisition by SpaceX (25:37)
As of November 12, OpenAI stops providing models to Cursor, citing compliance issues and an admission that xAI violated OpenAI’s terms of service. Both of them arrived at the same guess independently and neither claims it is in the article: if the accusation against the Chinese labs is that they distil American models to catch up, the same move is available to anyone who is behind. The other thing worth sitting with is that buying Cursor made SpaceX roughly the number three AI company in the US, with Google somewhere in fourth or fifth depending on how you count. Expect more consolidation as everyone scrambles to lock up resources.
-
Claude Fable and Mythos 5.1 (27:47)
Same model, two safeguard levels, with Mythos the permissive one for vetted cyber and life sciences work at US organizations only. 25% cheaper than Fable 5, 45% cheaper on agentic work, cache reads down to $0.25 per million. The headline results are protein binders roughly 10x better than the competition and an elevation map of Venus pulled out of decades-old radar data. Brett has not tried it and is deliberately staying on Opus 5 to save tokens; Travers has and says it is genuinely better on research problems. The segment turns into a conversation about flat-rate plans, which both of them are now completely dependent on, and their shared suspicion that no vendor can pull the plug first without their users walking across the street the same afternoon.
-
GLM 5.3 Flash will likely handle 45% of your AI workloads (34:01)
$0.15 per million input and $0.50 per million output, with a discount running into September. On the artificial analysis chart, Fable 5.1 Max with fallback sits at 66 on the intelligence index at about $4 a task, and GLM 5.3 Flash sits at 57 for roughly nine cents. GPT-5.6 Sol is 59 at 67 cents, Grok 4.6 is 61 at 94 cents. Travers makes the point the charts do not: you never see the error bars, only a mean, so the top few models are probably closer than four points suggests and the cost gap is not close at all. Their read is that anything rote, anything you would normally invoke a skill for, is a fine fit for the cheap model, and you save the expensive one for landing a real PR across a repo. The article also cites a McKinsey figure they both chewed on: 32% of respondents skipped a software purchase because of AI tooling, which is one in three people building instead of buying.
-
Serving GLM 5.3 Flash with vLLM (40:09)
320B total parameters, 18B active MoE, native FP8, 1M context, thinking always on, and one line to serve it. The 18B active number makes it sound local-friendly right up until you hit the roughly 306 GiB of weights. This is not a homelab job. You would want a heavily specced Mac Studio with well over 128GB of unified memory, and probably more than one. Which led straight into Brett browsing the Apple refurb store mid-episode and finding exactly two Mac Studios: one at about $10,000 running an M3, several silicon generations old, and the next one over $17,000. Computers have stopped depreciating, which is not a sentence either of them expected to say.
-
MiniMax H3 (44:09)
One model for text, image, video and audio. 15-second clips at 2K with native stereo, under a third the price of the mainstream options, open weights promised in the coming days, which is the same phrasing DeepSeek used a couple of weeks ago. Brett has no real use for image or video generation but wants to poke at it if it shows up in OpenCode, where MiniMax models usually land quickly. Travers rates the previous generation as the old king of price-performance before he moved to DeepSeek and GLM. On the same chart, Kimi K3 sits close to Fable 5 and neck and neck with GLM 5.3 Max, at about a dollar a task rather than four.
-
Build with Gemini Omni 1.1 Flash (46:30)
Studio-quality video with scene extension: feed it video and it extends in 10-second increments up to 40 seconds, filling in what it thinks should happen next, which is either cool or unsettling depending on the day. Drafts run at 360p, standard is 720p, output goes to 4K. Brett used the previous Omni for presentation video and images and says the quality was genuinely good, but in true Google fashion it took him longer to work out how to use the tool than to generate what he needed. He could not create a new API key at all, the button was greyed out with no explanation, so he reused an existing one. His verdict on paying Google money: here, take my cash, and it is like pulling teeth.
-
Gemini 3.5 Transcribe (50:44)
A new speech-to-text model that strips filler words and self-corrections, at 2.6% word error rate on pre-recorded audio, 4.0% streaming, across 85+ languages. The feature Brett actually cares about is intent resolution: “let’s meet Tuesday, no, Wednesday” comes out as Wednesday. That is a real complaint he heard from a customer about Google Meet transcription, where a correction later in a meeting leaves the transcript holding both answers and no way to tell which one won. He plans to run an episode of the show through it. The pitch about capturing natural speaking style to better understand intent is vague enough that they had to guess at it, landing somewhere between dialect and the way a question rises at the end of a sentence.
-
Model Hardware Standard research preview (54:17)
Anthropic’s attempt at a standard interface between models and physical hardware, so agents can safely operate devices. It uses a standardized driver, software that translates between an OS and hardware, extended to tell an agent how to use the device. The stated use case is cutting the time it takes to set up a lab environment: a library that hooks into Claude and drives a microscope or a robotic arm. Research preview, so nothing to run yet. Their reaction was that we are careening into science fiction, followed immediately by a long detour about getting Claude on a toaster, subscription toast, and laying bread directly on a GPU.
-
Agent skills that co-evolve with a persistent wiki (55:57)
A paper on co-evolving agent skills alongside a persistent knowledge base, reporting roughly a 10% bump in agent performance over time and beating existing skill-evolution methods. The mechanism is exactly what both of them have been building by hand: keep a wiki of skills and everything needed to run them, and iterate it as you learn more about the process you are automating. Worth reading if you care about agentic system design. Brett’s note on it was that it sounds great in theory, right up until the harness decides it does not need your skills today.
-
pstack: a skillset for rigorous agent engineering (58:09)
About 52 skills, with a claimed 2,000 PRs a month to production, and Claude Code ports already on GitHub. The core idea is that the thing worth maintaining is not the skills, it is the verification system: define bounds for the task, then keep reinforcing the verification loop until you can trust the output. There is a setup in there called Dr Eggbot that builds other bots, which they agree is a generational banger of a name. Brett wants to mine it against his own skills, particularly the skillify idea he built a few weeks ago, which worked well right up until the harness stopped loading skills.
-
Claude, Codex and Hermes installed unowned code inside corporate networks (59:59)
227 install commands across 120 corporate domains, pointing at packages nobody had registered, sitting inside llms.txt files. Those files are the new robots.txt for agents, and the errors are likely there because an agent generated the file in the first place. Agents doing ordinary research read them, treated vendor documentation as ground truth, and ran the install commands. Researchers registered a few of the unowned package names, stood up beacons, and had phone-homes from very large companies within hours. The conclusion is unpleasant and unavoidable: you have to treat any text an agent finds on the web as hostile, which means a sandbox layer between research and action. Brett’s plan is to put an llms.txt on his own site whose only instruction is to sign up for the newsletter.
-
Claude has become exhausting to use (62:52)
Brett found this while dunking on Claude all week and said “true dat,” which he would like you to know is how the cool people talk. Skills not running, and having to read War and Peace on every response. He asked for a one-word yes or no today, did not get one, asked why not, and got told he was right. The reason this matters beyond the griping is that reading what these things say to you is not optional, because the alternative is yoloing it and hoping. When every answer is a wall of text, the reading gets skipped, and the failure mode at the end of that road is an
rm -rfyou did not catch. Have backups. 3-2-1.
Off-the-Clock Recommendations
- Clue (1985) — Watched for movie night after Tim Curry’s death. It holds up, and it has one of the best closing lines in any film.
- Oblivion (2013) — Not this week’s pick, but it came up: Brett quotes “I don’t think we’re an effective team anymore” at Claude roughly once a day now.
- BirdNET-Go via security cameras — Security cameras already have microphones. Pipe the RTSP feed into BirdNET-Go in Docker on an old laptop or a NUC, run local models covering 6,000 species with BirdNET or 14,795 with Google’s Perch v2, and push into Home Assistant over MQTT. The author logged 418,726 detections and 271 species in 12 months. There is a bat-detection extension too. Brett is on vacation next week and setting this up, with a docker-compose file he is going to trust because it is on the internet, and it does not have an llms.txt.