RAHUL BHATIA / APPLIED AI & ML

Research depth.
Production scale.
Business impact.

I turn ambitious AI into systems that deliver.
From the first research question to hundreds of millions of daily requests.

Senior Applied Research Scientist at ServiceNow

6+ years translating science into valueTHE WORK SPEAKS
BUILT & SHIPPED AT
servicenowFidelity INVESTMENTSCREDRakuten

01 / THE TRACK RECORD

The scale is real.
So is the impact.

Enterprise AI. Financial services. Consumer technology.
A consistent focus on outcomes that matter.

200M+

Daily reranking requests

Production AI at ServiceNow, with sub-200ms p99 reranking latency.

SCALE / SERVICENOW
~250×

Lower summarization cost

GPT-4-level quality at roughly 1/250th the cost. Built for 20M+ calls a year.

EFFICIENCY / FIDELITY
32%

Campaign revenue lift

Geo-personalization that also reduced customer acquisition cost by 25%.

GROWTH / RAKUTEN BUSINESS-UNIT CAMPAIGN

02 / SELECTED WORK

From complex problems
to consequential systems.

A closer look at the work, the decisions, and the difference they made.

01 — SERVICENOW / ENTERPRISE AI

Making enterprise knowledge
actually useful.

Search that reasons across sources. Retrieval that understands images and tables. Infrastructure that keeps up with enterprise demand.

+31%Multi-hop answer accuracy
−26%Irrelevant results

Compared with a single-retriever baseline.

Multi-agent AIMultimodal RAGTriton
THE SYSTEM, SIMPLIFIEDRETRIEVAL → REASONING
Enterprise question
Planner agentDecompose · route · coordinate
Dense searchBM25 + metadataVisual retrieval
Synthesis & source arbitrationConfidence weighting · contradiction detection
MULTIPLE SOURCES. ONE GROUNDED ANSWER.
Explore the technical decisions

My ownership

Architected and led development of a multi-agent search pipeline: a Planner, specialized Retriever agents, and a Synthesizer, with A2A protocols and shared context.

Retrieval beyond text

Led multimodal document Q&A using ColPali visual retrieval and VLMs. Designed a tiling extension that improved retrieval performance 17% over the out-of-the-box baseline, with quantization and PCA compression to reduce patch-vector storage and latency.

Built for the production workload

Directed deployment of embedding, reranker, and multimodal models on NVIDIA Triton with dynamic batching, TensorRT, and FP16 quantization. Sustained 20M+ daily embedding requests at sub-50ms p99, alongside 200M+ daily reranking requests at sub-200ms p99.

Better training signals

Built dynamic hard-negative mining with teacher filtering, curriculum scheduling, and per-epoch refresh. Bi-encoder fine-tuning achieved 87% hit@5 and 76% recall@5 on an in-house multilingual corpus.

02 — FIDELITY / GENERATIVE AI
GPT-4 reference100%
Fine-tuned system~0.4%
RELATIVE SUMMARIZATION COST

Frontier-quality output.
Practical economics.

Made large-scale call summarization economical and turned a year of customer interactions into a briefing support teams could use.

LLM fine-tuningCustomerDNAGPU optimization
Read the story

The product

Conceived and led CustomerDNA: a generative AI product synthesizing a customer’s full-year interaction history across five channels into a single briefing, adopted enterprise-wide.

The engineering

Spearheaded a call summarization system producing 2–3 line summaries at GPT-4-level quality for roughly 1/250th the cost. Scaled to 20M+ calls annually on a four-node GPU cluster with quantization.

Additional research & leadership

Built GraphRAG for S&P 500 earnings reports and market news; devised CRAFT, a cost-reduced transformer fine-tuning architecture; developed few-shot classification across 500 call intents; and led interns fine-tuning Stable Diffusion for marketing banners, outperforming DALL-E-2 on in-house evaluation metrics.

03 — CRED / APPLIED MACHINE LEARNING
$1.34MAPPROXIMATE MONTHLY SAVINGS

Better data.
A stronger lending business.

Address intelligence that improved collections efficiency, expanded the whitelisting funnel, and reduced wasted field visits.

Representation learningRankingClustering
Read the story

The system

Built an address ranking and enrichment pipeline ingesting 15,000 new addresses a day from SMS data. Lifted collections efficiency 10% and the whitelisting funnel 4–5%, saving approximately $1.34M per month for the lending business.

The supporting models

Identified undeliverable addresses with 96% accuracy, improving collections operational efficiency by 5%. Built character-embedding deduplication with hierarchical clustering at 94% accuracy to eliminate duplicate field visits.

04 — RAKUTEN / PERSONALIZATION

Turning movement
into customer understanding.

Terabyte-scale GPS trajectories became Geo User Personas, better targeting, and measurable campaign growth.

Read the story

Built geo-personalization models that drove a 32% revenue increase and 25% CAC reduction in a business-unit campaign. Developed nearest-POI and transport-mode classification models processing a monthly batch of 300M points across 100K users in under five minutes with PySpark.

Built encoder-decoder trajectory embeddings and a purchase-probability classifier with 0.87 AUC to predict a retail chain’s next offline transactions.

300Mlocation points processedin under 5 minutesMONTHLY BATCH · PYSPARK

Results reflect individual projects and their stated evaluation contexts. Commercial systems are described at a high level; proprietary code and data are not shared.

03 / HOW I WORK

Scientific rigor.
Engineering judgment.

I connect the research question to the production constraint
and the business reason for solving it.

01

Frame the right problem.

Start with the decision a system needs to improve. Define the baseline, the success measure, and the constraints before choosing the model.

Applied researchEvaluationExperimentation
02

Build for the real world.

Accuracy matters alongside latency, cost, throughput, and maintainability. Design the system around the workload it must actually serve.

Distributed inferenceOptimizationProduction ML
03

Own the outcome.

Carry the idea through implementation and adoption. Lead technical decisions, mentor emerging talent, and keep the result tied to business value.

Technical leadershipMentorshipProduct thinking

04 / THE JOURNEY

Increasing scope.
Consistent ownership.

SEP 2024 — PRESENTNOW

ServiceNow

Senior Applied Research Scientist

Enterprise search, multimodal retrieval, and high-throughput AI infrastructure.Hyderabad, India

DEC 2022 — SEP 2024

Fidelity Investments

Data Scientist, L5

Generative AI for client servicing, efficient LLMs, and financial knowledge systems.Bengaluru, India

NOV 2021 — DEC 2022

CRED

Data Scientist

Address intelligence and applied ML for lending and collections.Bengaluru, India

JAN 2020 — NOV 2021

Rakuten

Associate Data Scientist

Geo-personalization, spatio-temporal modeling, and customer representation learning.Bengaluru, India

MAY 2019 — JUL 2019

InnovAccer

Data Science Intern

Interpretable patient no-show prediction: 0.85 AUC-ROC and an estimated $50K in savings for a business client.Noida, India

THE FOUNDATION

B.Tech · Computer Science
The LNM Institute of Information Technology, Jaipur

2016 — 2020

05 / BEYOND THE DAY JOB

Build. Publish.
Give back.

RESEARCH / ACL 2026 · ARGMINING WORKSHOP

RESOLVENOW

Rule-Based Type Classification with LLM-Driven Multi-Label Tagging for UN Resolutions.

Vedant Gupta · Rahul Bhatia
Vaibhav Varshney · Manjunatha Naik

View on ACL Anthology ↗
SHARING THE CRAFT

International conference speaker.

Interpretable ML and categorical encoding at PyCon Malaysia 2019 and PyCon Indonesia 2020.

PyCon speaker listing ↗
OPEN SOURCE & MENTORSHIP

Helping the next generation build.

Google Summer of Code 2019 student mentor. Mentored two students as a long-time Public Lab open-source contributor.

Explore my GitHub ↗
ENTREPRENEURIAL PERSPECTIVE

Silicon Valley Fellow.

Selected for a state-sponsored startup exposure trip with a 0.39% acceptance rate. Pitched an AI product to investors at Silicon Valley Bank, California.

THE EARLY EXPLORATIONS

A builder before the job title.

Public projects and writing across ML, finance, healthcare, and forecasting.

Read my technical writing on Medium ↗
Rahul BhatiaRAHUL BHATIA / HYDERABAD, INDIA

THE PERSON BEHIND THE SYSTEMS

A researcher’s curiosity.
A builder’s instinct.

My work has taken me from healthcare prediction and consumer personalization to generative AI and enterprise search. Across those domains, the challenge I enjoy most stays the same: turning a difficult, ambiguous problem into something useful.

I work across research, engineering, and product. I’m equally invested in why a model works, how it behaves under production constraints, and whether it makes a meaningful difference to the people using it.

Generative & agentic AISearch & recommendationsMultimodal learningProduction ML systems
My technical toolkit

Languages: Python, Scala, SQL, Java

Models & methods: LLMs, RAG, GraphRAG, agents, ColPali, VLMs, embeddings, transformers, diffusion models, contrastive learning, fine-tuning, evaluation, recommendation systems

Systems: NVIDIA Triton, TensorRT, quantization, distributed training and inference, GPU clusters, PySpark, Hive, Snorkel

OPEN TO THE RIGHT OPPORTUNITYRESEARCH · ENGINEERING · TECHNICAL LEADERSHIP

Your next ambitious
AI problem.
Let’s make it real.

Exploring applied science, data science, ML engineering,
and AI leadership opportunities.