Now
What I’m working on, reading, and thinking about right now.
Updated 5 August 2026
A now page — what’s currently filling my head, in 60 seconds.
Working on
- GPU inference at scale — tuning open-weight model serving on Azure: vLLM, tensor parallelism, and the KV-cache pressure that decides your real concurrency ceiling.
- HPC benchmarking for scientific workloads — profiling long, multi-stage pipelines on H100 clusters where the bottleneck is almost never where the vendor chart says it is.
- Helping enterprise teams get from “we have GPU quota” to a serving stack they can page someone about at 3 a.m.
- KCSA next — KCNA passed, four to go on the Kubestronaut run. Study guides go up as I write them.
Reading
- A stack of recent inference-systems papers — paged attention, speculative decoding, disaggregated prefill/decode. Notes accumulating.
- The Coming Wave — Mustafa Suleyman. On a second pass. Full review →
- Alignment evals papers — write-ups landing in /ai-safety.
Building
- pi-bench — grading whole prompt-injection defense stacks on ASR, false positives, latency, and cost rather than testing detectors one at a time.
- oss-model-playbook — the deployment playbook for open-weight LLMs on on-prem K8s, AKS, and Azure AI Foundry.
- vector-engineering-for-agents — the retrieval layer, at the depth needed to defend it in a design review.
- This site. From zero, in public, with Hugo.
Thinking about
- Why served p99 and benchmark throughput diverge, and how much of the difference is scheduling rather than silicon.
- Whether scalable oversight can outpace capability scaling. (Short answer: only if oversight itself is automated.)
- Why every enterprise AI rollout I’ve seen succeeds or fails on the boring layer, not the model.
Outside work
- Cycling around the Zürichsee — back into base-mile mode.
- Zürich → Berlin on the train often enough that I have favorite carriages.
Inspired by Derek Sivers. If you have a now page, send it to me.