Open to Opportunities

Manideep Reddy Mandhadi

Machine Learning Engineer — Agentic AI, LLM Evaluation & MLOps

Sunnyvale, CA

Machine Learning Engineer with 5+ years building multi-agent AI, LLM evaluation systems, and ML infrastructure across financial services and enterprise domains. Specializes in RAG and agentic workflows, evaluation frameworks (LLM-as-a-judge, rubric scoring, GEPA), model fine-tuning (LoRA/QLoRA/PEFT, RLHF/DPO), and experimentation & causal inference. Proven at turning research into observable, continuously-improving production systems with robust MLOps on AWS SageMaker, Docker, and Kubernetes.

Technical Skills

A deep toolkit spanning agentic AI, model alignment, experimentation, and production MLOps.

LLMs, Agents & Evaluation

Multi-agent orchestration Tool-calling Guardrails & fallback MCP servers RAG Graph RAG Vector search Reranking Prompt engineering LLM-as-a-judge Rubric scoring GEPA optimization Drift detection Persona simulations Human-in-the-loop

Model Adaptation & Alignment

Fine-tuning LoRA QLoRA PEFT SFT RLHF RLAIF DPO GRPO Reward modeling Embeddings

Experimentation & Causal Inference

A/B testing Switchback experiments Geo-experiments Lift / incrementality Uplift modeling Bayesian methods Causal inference Hypothesis testing

LLMOps & Observability

LangChain LangGraph Langfuse Hugging Face Tracing Evaluation pipelines Experiment tracking

MLOps & Infrastructure

AWS SageMaker Model serving REST APIs Model registry Feature store Drift monitoring MLflow Kubeflow CI/CD Docker Kubernetes

Deep Learning & NLP

PyTorch TensorFlow Transformers BERT LSTM / RNN CNNs NER Summarization Time-series forecasting Speech-to-text

Computer Vision

Object detection Image segmentation OCR Autoencoders GANs / diffusion OpenCV

Data & Cloud

AWS S3 Lambda SQS / Kinesis DynamoDB Rekognition Textract Databricks Snowflake Spark Pandas Polars SQL

Vector / Graph & Search

Pinecone Qdrant Milvus ChromaDB OpenSearch Neo4j

Programming

Python TypeScript JavaScript SQL

Work Experience

Turning research into observable, continuously-improving production systems.

Machine Learning Engineer

Intuit
Jul 2025 – Present Mountain View, CA
  • Built and productionized multi-agent AI for TurboTax and QuickBooks financial workflows, released first to internal experts and sales teams, then to customer-facing products.
  • Built RAG retrieval pipelines using vector search, chunking, and reranking, and orchestrated multi-agent workflows with tool-calling, guardrails, and fallback, including MCP servers for agent-to-tool integration.
  • Designed and analyzed A/B tests, switchback, and geo-experiments to measure lift/incrementality, applying Bayesian methods, uplift modeling, and causal inference.
  • Designed an LLM-as-a-judge evaluation framework with rubric criteria (correctness, no-hallucination, completeness, relevance, compliance, tone), tracking false positives/negatives and improving judge alignment, cutting manual review effort by ~40%.
  • Built a self-improving evaluation harness applying GEPA-style optimization over a Pareto frontier with train/validate/test splits, auto-tuning agents against ground truth.
  • Instrumented end-to-end tracing (Langfuse) to classify system behavior and cluster failures into themes for error analysis; curated golden datasets via human-in-the-loop smart labeling.
  • Built drift detection and hierarchical, taxonomy-based clustering for topic discovery; ran persona-driven simulations to stress-test agents pre-launch.
  • Developed document-intelligence models (classification, OCR, entity extraction) for millions of financial documents at >90% classification and ~93% extraction accuracy.
  • Delivered capabilities that reduced handling time per case by ~30% and saved experts ~5 hours/week.
  • Evaluated and compared large foundation and vision-language models (Gemini, Claude, Nova, Llama) by accuracy, latency, and cost.

Machine Learning Engineer

Toyota Motor Corporation
Jun 2023 – Jul 2025 Los Altos, CA
  • Fine-tuned and adapted foundation models with LoRA/QLoRA/PEFT/SFT to build internal tools for developers and sales teams.
  • Applied RLHF/RLAIF/DPO and reward modeling to align assistant responses with expert and user feedback.
  • Built RAG retrieval pipelines integrated with Pinecone, Neo4j, and OpenSearch to power internal knowledge assistants.
  • Fine-tuned BERT and created LangChain-based assistants for domain-specific Q&A with guardrails and fallback mechanisms.
  • Deployed real-time deep learning models (YOLOv7 + TensorRT) to edge devices via Docker/Kubernetes, cutting inference latency to <20ms.
  • Led predictive-maintenance pipelines for EV-battery diagnostics using LSTM forecasting and Isolation Forest over AWS S3/Snowflake/Polars.
  • Built a driver-behavior analytics system (K-Means, DBSCAN, Random Forest) on Databricks across large-scale telematics data.
  • Automated end-to-end ML workflows with MLflow, Kubeflow, and GitHub Actions CI/CD.
  • Designed CNN, RNN, and Transformer architectures spanning vision, NLP, and time-series tasks.
  • Instrumented model interpretability (SHAP, LIME) and deployed models on AWS SageMaker and Azure ML.

Machine Learning Engineer

Siemens
Jan 2021 – Jul 2022 India
  • Built autoencoder- and Isolation-Forest-based streaming anomaly detection on the industrial IoT platform (Kafka/Spark), enabling real-time alerts.
  • Developed LSTM time-series predictive-maintenance models for PLC-controlled equipment, reducing unplanned downtime by 22%.
  • Wrote REST APIs with FastAPI and containerized them using Docker, allowing industrial applications and dashboards to access models instantly.
  • Built end-to-end ML pipelines (Scikit-learn, Pandas, MLflow) for experiment tracking and deployment readiness.
  • Trained and deployed deep learning classification/regression models (TensorFlow, Keras, PyTorch) for batch and real-time serving.
  • Established ML CI/CD with Docker, Git, Jenkins, and FastAPI for reproducible packaging and automated deployment.
  • Ran model evaluation, hypothesis testing, and A/B testing to validate performance before rollout.
  • Created React.js/Plotly dashboards enabling operations teams to monitor anomaly alerts and predictive scores in real time.
  • Prototyped LLM automation (OpenAI, LangChain) for document classification and summarization within the AI CoE.

Credentials

Certifications and academic foundation in data science and engineering.

Certifications

AWS Machine Learning – Specialty

Advanced ML on AWS

AWS Data Engineer – Associate

Data pipelines & architecture

Education

University of the Pacific

M.S., Data Science — San Francisco, USA

JNTUH

B.Tech, Electronics & Communication Engineering — India

Get in Touch

Interested in collaborating or have an opportunity in mind? Let's talk.

manideepreddy965@gmail.com (415) 902-1124