Model Releases & Capabilities
Google AI Blog
Autonomous LLM post-training with Tunix on TPUs
Google published a piece on Tunix, their JAX-native library for post-training LLMs on TPUs, framed around a scenario where you write a single Markdown spec and an agent handles the RL fine-tuning loop autonomously overnight. It's part of the broader push to make post-training (especially RL-based fine-tuning) something you can hand off to an agent rather than babysit.
Why it matters
I'm intrigued but wary of the framing — 'write a spec, wake up to a trained model' is the kind of pitch that undersells how much eval and guardrail work still has to happen before you'd trust the output. Still, TPU-native tooling for RL post-training is an area I don't think gets enough attention compared to pretraining, so I'll be watching where this goes.
Tools, Infrastructure & Open Source
Claude Code Changelog
v2.1.268
v2.1.268 adds enterprise gateway features: a `pricing:` setting in `gateway.yaml` so signed-in Claude Code clients get consistent rates through managed settings (keeping `/cost` and telemetry in sync with the spend meter), a startup warning when a gateway's `access_control.allow_cidrs` is empty, a one-time warning on the first request from a public address, and a `gatewayInternalNetworks` managed setting for allowing internal `/login` access.
Why it matters
This reads like Anthropic hardening the self-hosted gateway path for larger orgs — the empty-CIDR warning in particular looks aimed at catching a real misconfiguration people were probably hitting. Not flashy, but exactly the kind of guardrail I'd want if I were running one of these gateways myself.
Claude Code Changelog
v2.1.269
Claude Code v2.1.269 adds `claude plugin eval` for scored, reproducible eval runs against a plugin (JSON + HTML report), a `/output-style` command to list and switch output styles — including over Remote Control and headless sessions — a diff of files a Bash command changed surfaced in the Bash tool result, and an `OTEL_METRICS_INCLUDE_REPOSITORY` flag for tagging telemetry by repo.
Why it matters
The plugin eval command is the one I care about most — it turns 'does my plugin actually work' from a vibe check into something you can run in CI. The Bash-edit diff is a small but genuinely useful visibility win too; I've lost track of file changes from a Bash call more than once.