Data Scientist · Experimentation, Causal Inference & Applied ML

Pushpendra Singh

I make A/B tests conclusive and recommendations relevant.

Senior Consultant at Straive, working across 64 A/B experiments and 17.6M visitors for a large-scale e-commerce client. I designed and built Jetstream — an AI-powered experimentation platform — and the retrieval layer behind a personalized homepage carousel.

  • Python
  • SQL
  • A/B testing
  • CUPED
  • XGBoost
  • Tableau
  • FastAPI
  • Senior Consultant, Straive
  • IIT (BHU) Varanasi
  • ex-EXL
  • Hyderabad, India
Pushpendra Singh, Data Science & AI Consultant

Measured impact

$6.8B

catalog served by the item-similarity engine I built end-to-end

0.88

ranking AUC against a 0.49 baseline, retrieving over 888K products

$37M

revenue attributed in a single campaign analysis, +46.5% YoY

9.1%

median standard-error cut from a variance-reduction technique I designed

Profile

A numbers-first read on what actually moves the business

I'm a data science and AI consultant based in Hyderabad, currently a Senior Consultant at Straive, where I build machine learning systems and experimentation infrastructure for a large-scale e-commerce client. My work spans causal inference and variance reduction for A/B testing, recommendation and candidate-generation systems, and Jetstream, the internal experimentation platform I designed and built end to end.

Earlier, at EXL, I built and maintained 100+ Tableau dashboards used by senior leadership and ran campaign analytics on budgets exceeding $800K/month, attributing tens of millions of dollars in revenue across direct mail, email, and paid search.

I hold an Integrated B.Tech + M.Tech from IIT (BHU) Varanasi, and taught myself the ML/AI stack that now anchors my work — from classical statistics and experiment design through modern LLM application development.

Experience

Feb 2026 — Present

Senior Consultant

Straive · Hyderabad · Data Science & AI

Current
  • Designed and built Jetstream, an AI-powered experimentation platform where A/B tests are planned, monitored, and decided in one place — a production CUPED variance-reduction engine, chi-square Sample Ratio Mismatch checks that can block a bad ship decision, and a daily AI briefing across the whole experiment portfolio. Built solo end to end: FastAPI + React, ~10,000 lines, 64 unit tests on the statistics engines, containerized and deployed as a feedback beta.
  • Rebuilt the conversion covariate behind app experimentation, benchmarked eight variance-reduction estimators, and designed two original refinements — mechanism-matched metrics and a "Zero-to-Hero" noise-reduction technique that cut standard error by a median 9.1% across the 36 of 64 experiments where it applied, roughly 3× the ceiling of covariate tuning alone — taking conclusive experiments from 13 → 24 of 64 at full test coverage.
  • Built the retrieval layer behind New For You, a personalized homepage carousel — clustering each shopper's history into interest centroids and retrieving against an 888K-product embedding space. Lifted ranking AUC to 0.88 from 0.49 against the user-embedding baseline, took shoppers scoring above chance from 47% → 94%, and nearly doubled catalog coverage to 57%.
  • Shipped an executive out-of-stock dashboard quantifying $74.7M in annual recoverable revenue across the client's North America e-commerce, engineering a two-segment methodology that combines behavioral session data with inventory time-series modeling to deliver CFO-ready numbers.
  • Built an end-to-end item-item similarity engine for 221K+ products, blending 1,024-dimensional semantic embeddings with structured attribute scoring to surface the ten most relevant items per product across a catalog driving $6.8B in sales.
  • Ran persona-level demographic analysis across 2M+ customers, segmenting five behavioral personas to surface $234M in revenue patterns and inform merchandising strategy.
Read more on the technical approachShow less
  • Jetstream — design discipline built into the product. Setup is a stepped flow ending in an "Experiment Contract" that records the pre-specified primary analysis method, locked at launch. The decision engine only uses the CUPED-adjusted result when CUPED was pre-specified beforehand and the covariate is eligible, and states which basis it used on screen — the guardrail that stops post-hoc method shopping. Standard and CUPED results are reconciled so no two screens can disagree: the point estimate never moves, only the standard error shrinks, so CUPED changes your confidence in a result and never its direction.
  • Jetstream — AI that can't invent things. The daily monitoring pass scans the portfolio and diffs against the previous briefing, but the scan itself is deterministic: the model narrates it and cannot add or drop items, so briefings are reproducible and auditable. Generated insights must cite the evidence they came from — uncited insights are dropped rather than displayed. Anomaly detection runs statistically on daily per-variant series before the model is asked to explain a cause.
  • New For You — diagnosis and honest evaluation. I diagnosed why the baseline failed rather than just beating it: user–product cosine similarity sat near 0.2 against roughly 0.9 for interest–product, meaning the user embedding wasn't living in the same space as the product embeddings. I then validated against what shoppers actually did next across 77,439 shoppers and 384,313 next-actions, which showed 75% of next-actions were re-engagement with already-seen products that the product excludes by design — so the novelty filter, not the model, caps the achievable recall. Recall@300 of 12.2% sits close to the ~16% structural ceiling, and I designed the evaluation with a historical cutoff to rule out leakage.
  • Variance reduction — the finding that mattered more than the gain. I added 56 app-native behavioral features to the existing 64 and cross-fit predictions by experiment so no visitor's score came from a model that had seen their own test, avoiding the leakage that would bias CUPED. That covariate work hit a hard ceiling — correlation with conversion topped out near 0.24 and ROC-AUC near 0.68 across nine model families — so I proved the ceiling rather than keep chasing it, and moved the leverage elsewhere. Mechanism-matched metrics (judging each test on the metric its change actually targets) took detections from 13 → 20 of 64; "Zero-to-Hero," which imposes a known-zero effect on visitors who couldn't have been affected instead of estimating it, added a median 9.1% standard-error cut and brought the final count to 24. A looser 80% confidence bar would have shown 28, but roughly 7 of those are expected false positives, so I rejected it. The headline conclusion I delivered was that precision was never the bottleneck: the median non-significant experiment sat ~17× of traffic away from significance, so effect size and traffic were the real constraint — and I recommended bigger treatments and more sensitive primary metrics over more statistics.
SE reduction — production methods 1.8% SE reduction — rebuilt ML covariate 3.0% SE reduction — Zero-to-Hero (my technique) 9.1% median Conclusive experiments of 64 13 → 24

Across all 64 evaluated experiments at 100% coverage; production variance-reduction methods covered under 50%. Zero-to-Hero applied to the 36 experiments with an eligible pre-exposure signal.

Oct 2023 — Feb 2026

Lead Assistant Manager (Consultant II)

EXL Services · Gurugram · Marketing Analytics · 2 years 5 months

  • Promoted from Consultant to Consultant II within 1.5 years, and recognized with the Stellar Performance Award for high-impact delivery and client value.
  • Drove "Offer Not Taken" email campaign analysis attributing $37M in revenue at 46.5% YoY growth, with a real-time Tableau dashboard for ongoing tracking.
  • Led Direct Mail campaign analytics informing a $100K investment projected to generate $5.6M+ in revenue over two years.
  • Built a target-ranking model for sales prioritization and automated reporting pipelines; maintained 100+ Tableau dashboards across paid-search and direct-mail programs with budgets over $800K/month.

Earlier

Zestech
Associate Technical Trainer — trained 300+ students in Python
2023
Qolaba
ML Engineer — LangChain and OpenAI chatbot on a Pinecone vector database
2022 — 2023
DeepEdge
ML Engineer — dataset-quality visualization pipelines and ONNX model standardization
2022
Clean Electric
Structural Engineer — battery-pack casing and FEM simulation
2019

Capabilities

Experimentation & causal inference
A/B testingCUPEDCUPACANCOVAVariance reductionCross-fittingCausal inferencePower analysisExperiment design
Machine learning & AI
XGBoostLightGBMEmbeddingsSemantic searchRecommender systemsClaudeOpenAILangChainOptunaONNX
Analytics & BI
SQLTableauCampaign analyticsDashboardingDemand & inventory analysisStakeholder reporting
Engineering
PythonFastAPIReactDockerGitPipeline automationData modelingUnit testing

Get in touch

Let's discuss the role.

Happy to walk through the approach behind any of this work — the methods, the trade-offs, and what the results actually meant for the business.