Blog

Thoughts on building Multimodal AI, ML systems at scale, and lessons from the trenches.

16 min read

RLHF Post-Training: Designing the Training and Serving Systems

RLHF is usually explained as an objective. An explainer on the systems it implies: four model copies resident at once, an inference engine and a trainer in the same job, and a serving path where the thing you optimised against never ships.

explainerllmpost-trainingrlhfml-systemssystem-design
Read post →
7 min read

Safety Policy as Retrieved Context

Most safety behaviour is frozen into weights at training time, while the policy it encodes changes weekly. An explainer on retrieving the clauses that apply to a request and inlining them as explicit constraints, and the failure modes that introduces.

explainerai-safetyragllmguardrails
Read post →
18 min read

Designing a RAG Search System: Indexing, Retrieval and Grounding

Retrieval-augmented generation is usually drawn as three boxes. An explainer on the parts that decide whether it works: chunking and reindexing, hybrid retrieval and reranking, permissions enforced at retrieval, and the difference between a bad answer and a bad retrieval.

explainerragsearchretrievalllmsystem-design
Read post →
8 min read

Transformer Recommenders: Designing the Training and Serving Systems

Sequential transformer recommenders changed what the model is. An explainer on the systems work that decides whether that change survives contact with production: point-in-time correctness, the retrieval/ranking funnel, and the skew that quietly eats the gains.

explainerrecommendation-systemstransformersml-systemssystem-design
Read post →
5 min read

Can Geometry Predict Whether a Face Matches a Voice?

We show that the intrinsic Riemannian geometry of pretrained neural network embedding spaces predicts how well face and voice can be matched across modalities: without any cross-modal training.

multimodal-airiemannian-geometrybiometricscomputer-visionresearch
Read post →
5 min read

Do LLM Recommenders Obey Preference Axioms?

We test whether LLM-based recommender systems satisfy classical rationality axioms from social choice theory, and find that all models violate every axiom, with a striking dichotomy between pairwise and set-based reasoning.

large-language-modelsrecommender-systemssocial-choice-theorylogical-reasoningresearch
Read post →
6 min read

Are VLM Identity Judgments Logically Consistent?

We test whether vision-language models obey symmetry and transitivity when judging if two images show the same person, and find a striking accuracy-consistency trade-off.

vision-language-modelsperson-re-identificationlogical-reasoningcomputer-visionresearch
Read post →
3 min read

Building Multimodal AI Systems That Serve a Billion Recognitions a Day

Lessons learned from building and scaling multimodal person recognition at Amazon: fusing voice, face, Bluetooth, and behavioral signals to identify users across millions of devices.

multimodal-aiperson-recognitiondistributed-systemsmachine-learning
Read post →