Available for collaboration

Raj Thakur - Senior ML Engineer

Senior ML Engineer at Amazon AWS

Specializing in ML frameworks optimization and acceleration on AWS Trainium & Inferentia chips. Expert in PyTorch, JAX, and custom silicon optimization for high-performance ML systems.

About Raj Thakur - Senior ML Engineer

I'm a Senior Machine Learning Engineer at Amazon. My job is making large-model inference fast on AWS Trainium and Inferentia. I have 10+ years of experience building ML systems that serve and train models at cloud scale.


Day to day I work on LLM serving in vLLM, KV cache and decode-path optimization, PyTorch and JAX on custom silicon, and the PJRT runtime. Most of what I build ends up in the AWS Neuron SDK that customers use on Trainium and Inferentia.

ML Engineering Skills & Technical Expertise

🧠

ML Frameworks & Deep Learning

PyTorch XLA, JAX, PJRT Runtime, Transformers, CNNs, RNNs, LSTMs, Attention Mechanisms, BERT, GPT architectures

ML Acceleration & Custom Silicon

AWS Trainium, AWS Inferentia, Custom Silicon Optimization, Hardware-Software Co-design, Performance Profiling, Memory Optimization

🔄

Distributed ML & Training

Data Parallelism, Model Parallelism, Gradient Synchronization, Distributed PyTorch, Multi-node Training

🚀

ML Operations & Deployment

Model Serving, A/B Testing, ML Pipelines, Kubernetes, Docker, CI/CD for ML, Model Monitoring

📊

Data Engineering & Processing

AWS Services, S3, Lambda, Step Functions, ElasticSearch, OpenSearch, Apache Spark, Kafka, ETL Pipelines, Data Lakes, Real-time Streaming, Big Data Analytics

🔬

Programming Language

C++, JAVA, Python, Shell

Professional Experience - Amazon AWS & Tech Companies

Amazon Web Services (AWS)

6 yrs 3 mos • Full-time

Senior Machine Learning Engineer

Dec 2022 - Present • Calculating... Cupertino, California, United States

Making large-model inference fast on AWS Trainium and Inferentia, and building the native PyTorch device path for both

  • Built disaggregated inference for vLLM on Neuron, so prefill and decode run on separate Trainium nodes with the KV cache moved device to device to cut serving latency
  • Designed the runtime batch operations API for KV cache transfers, which doubled throughput, and tuned EFA utilization for inter-node tensor transfers
  • Got disaggregated inference ready for product launches, including multi-node configurations and mixed GPU plus Trainium deployments, validated on the EKS production stack
  • Won back per-step host time that asynchronous execution was supposed to hide, and found the real cause of a decode stall on a 64-way expert-parallel deployment after it had been blamed on prefill
  • Unblocked a large mixture-of-experts model that would not fit in device memory, and brought accuracy back in line with the XLA baseline
  • Made Trainium run stock PyTorch. Models compile through standard torch.compile instead of a forked PyTorch/XLA stack, so customers keep the programming model they already know
  • Built CPU-only compilation for torch.compile, which PyTorch itself does not support, so models compile without taking accelerator capacity
  • Brought up native tensor-parallel vLLM serving end to end on Trainium3
  • Built the PJRT backend for AWS Trainium that lets JAX run on custom silicon
  • Optimized the PyTorch XLA integration for 1.2x faster training on Trainium
  • Built HLO-based graph optimizations and drift detection tooling
vLLM LLM Serving KV Cache Optimization PyTorch JAX AWS Trainium AWS Inferentia

Software Development Engineer II

Jun 2020 - Nov 2022 • 2 yrs 6 mos Bengaluru, Karnataka, India

AWS Billing platform engineering and OpenSearch distributed systems

  • Built Monetization Authority from scratch, a Step Functions workflow engine that automates enterprise billing document generation across thousands of payers. It drove $300M in free cash flow improvements
  • Designed a transfer billing system that simplified billing for third-party sellers on AWS and improved third-party account security posture, estimated to add $1.2B in annual revenue by 2027
  • Built a payer classification system and automated payer mapping, which recovered another $100M in free cash flow
  • Led the JDK 17 migration across 8+ packages and cross-functional teams of 6+ developers
  • Containerized AWS OpenSearch data plane components and built an automated node diagnostic and self-healing framework that cut customer RCA tickets by 30%
AWS Step Functions AWS Lambda Redshift OpenSearch Distributed Systems Java

Grab

2 yrs • Full-time

Senior Software Engineer

Jul 2018 - Jun 2020 • 2 yrs Bengaluru Area, India

Lead Engineer in Settlement Platform for Grab Financial Group

  • Scaled the settlement platform from worker-based to event-driven architecture, improving merchant settlement performance by 60%
  • Designed the core post-payment processing unit for payment gateways and international merchants in Golang
  • Built payment gateway integration systems handling millions of transactions
Golang Fintech Payment Systems Event-Driven Architecture

Yatra Online Pvt Ltd

9 mos • Full-time

Software Engineer, Full Stack

Nov 2017 - Jul 2018 • 9 mos Hyderabad, Telangana, India

Full-stack development for corporate travel booking

  • Designed a customized Header Service for corporate merchants using Node.js, Java, and Spring, with HAProxy and Nginx routing
Node.js Java Spring Nginx

Tata Consultancy Services

2 yrs 3 mos • Full-time

Software Engineer

Aug 2015 - Oct 2017 • 2 yrs 3 mos Hyderabad, Telangana, India

Embedded systems test frameworks and enterprise software solutions

  • Developed the TCS-Ericsson JCAT framework for testing Ericsson AXE IO embedded systems
  • Built an AI-powered web console for test recommendations, recognized with a Star Performer award
  • Customized the OpenStack console UI for the ATLAS project
Java OpenStack Embedded Systems Full-Stack Development

Technical Writing

I regularly write about machine learning infrastructure, system design, and emerging AI technologies. My articles focus on practical insights from building production ML systems.

Editor of Software System Design publication on Medium, featuring in-depth articles on scalable architecture and distributed systems.

📱
UX Design

Split Bill App Design: A Guide to Creating a Seamless Splitwise-like Experience

Comprehensive guide to designing intuitive bill-splitting applications with focus on user experience.

Read Article →
🔧
Software Engineering

DI Frameworks: Spring, Guice and Dagger

Comprehensive comparison of dependency injection frameworks for large-scale applications.

Read Article →
🏗️
System Design

System Design Interview Primer

Essential tips and strategies for tackling high-level system design questions in technical interviews.

Read Article →

Let's Connect

Interested in discussing ML innovations, system architecture, or collaboration opportunities?