Blog
Thoughts on building Multimodal AI, ML systems at scale, and lessons from the trenches.
RLHF Post-Training: Designing the Training and Serving Systems
RLHF is usually explained as an objective. An explainer on the systems it implies: four model copies resident at once, an inference engine and a trainer in the same job, and a serving path where the thing you optimised against never ships.
Read post →Safety Policy as Retrieved Context
Most safety behaviour is frozen into weights at training time, while the policy it encodes changes weekly. An explainer on retrieving the clauses that apply to a request and inlining them as explicit constraints, and the failure modes that introduces.
Read post →Designing a RAG Search System: Indexing, Retrieval and Grounding
Retrieval-augmented generation is usually drawn as three boxes. An explainer on the parts that decide whether it works: chunking and reindexing, hybrid retrieval and reranking, permissions enforced at retrieval, and the difference between a bad answer and a bad retrieval.
Read post →Transformer Recommenders: Designing the Training and Serving Systems
Sequential transformer recommenders changed what the model is. An explainer on the systems work that decides whether that change survives contact with production: point-in-time correctness, the retrieval/ranking funnel, and the skew that quietly eats the gains.
Read post →Can Geometry Predict Whether a Face Matches a Voice?
We show that the intrinsic Riemannian geometry of pretrained neural network embedding spaces predicts how well face and voice can be matched across modalities: without any cross-modal training.
Read post →Do LLM Recommenders Obey Preference Axioms?
We test whether LLM-based recommender systems satisfy classical rationality axioms from social choice theory, and find that all models violate every axiom, with a striking dichotomy between pairwise and set-based reasoning.
Read post →Are VLM Identity Judgments Logically Consistent?
We test whether vision-language models obey symmetry and transitivity when judging if two images show the same person, and find a striking accuracy-consistency trade-off.
Read post →Building Multimodal AI Systems That Serve a Billion Recognitions a Day
Lessons learned from building and scaling multimodal person recognition at Amazon: fusing voice, face, Bluetooth, and behavioral signals to identify users across millions of devices.
Read post →