Anuj Dutt.
ML Tech Lead @ Adobe
I architect and ship production AI systems — from cloud-scale GenAI and GPU-optimized LLM inference to edge AI on constrained devices. I specialize in taking models from research to production through hardware–software co-design, aligning models, runtimes, and systems with real constraints in latency, memory, compute, and reliability.
At Adobe, I lead architecture and delivery for GenAI systems across Acrobat initiatives, including PDF Spaces, Contract Understanding and Generate Presentation. My work spans workflow architecture, multi-agent orchestration, evaluation, and production reliability. I set technical direction across engineering, product, design, and research—aligning stakeholders around key decisions and mentoring engineers from experimentation through launch.
Previously, as ML Tech Lead and AI Architect at Jabra, I led edge-AI delivery across three commercial video product lines. I shipped the on-device person detection, tracking and segmentation behind intelligent framing and background effects on PanaCast 50, PanaCast 50 VBS and PanaCast 20 — reducing model memory by 16% and improving performance by more than 20% within a fixed hardware envelope. I owned the SoC platform decision that unblocked a next-generation product architecture, and authored the core algorithm behind the first Intelligent Meeting Spaces release.
Earlier, in Bose's Consumer Electronics Applied Research group, I built on-device AI for the Bose AR ecosystem and delivered Bose's first mobile on-device neural network, shipped in the BoseAR library. I developed an IMU-only system that recognized ten gestures from head motion — without a camera or microphone.

0
Products in market
0
Granted US patents
0K+
GitHub stars
0K+
Technical video views
0
Official CVPR tutorials
Ideas & influence
Industry and community leadership.
I teach at conferences, publish open-source projects and technical education, and help practitioners solve machine-learning problems.
Latest Writing

Understanding LLM Context Length: What It Really Means and Why It Matters?
A deep dive into why context length exists in LLMs, what breaks when you exceed it, and the architectural constraints that make training long-context models expensive.

LLM in a Flash: Efficient LLM Inference with Limited Memory
Exploring techniques for running large language models efficiently on memory-constrained devices, optimizing inference without sacrificing performance.

Emerging Properties in Self-Supervised Vision Transformers: DINO Paper Summary
A comprehensive summary of the DINO paper exploring emerging properties in self-supervised vision transformers and their applications.
Weights & Wires
Deep dives on production AI — LLM inference, on-device models, and the systems around them. Roughly monthly.
Join 2,400+ technical readers
Start a conversation
Building an ambitious AI system?
I'm always interested in serious conversations about production GenAI, LLM inference, Edge AI, hardware–software co-design and technical leadership.