data & AI engineering

Aman Trivedi

Data and AI Engineer — ML, Scorecards & Data Platforms

I build statistically grounded ML models and interpretable scorecards behind credit and collections decisions, and the AWS data pipelines that feed them.

Data and AI Engineer at AYE Finance, a regulated NBFC, and M.Sc. (Data Science & AI) candidate at BITS Pilani. 2nd Runner-up, AWS Agentic AI Hackathon; HackerRank SQL 5★.

résumé ↗selected work ↓

4B+rows processed by owned pipelines
8+ETL / ELT pipelines owned
2h → 15mMIS report latency
3 monthsData Warehouse ahead of deadline
5,000+employees on the analytics platform

Selected technical initiatives

01

XGBoost · ranking · pilot

Collections Prioritization model

Ranks overdue accounts by recovery likelihood, paired with an LLM-generated reporting layer that narrates segment trends. Live for a pilot group in real-user testing ahead of wider rollout.

  • XGBoost
  • Ranking
  • LLM reporting

02

WOE/IV · imbalance · monitoring

Interpretable credit scorecard

A Logistic Regression scorecard with WOE/IV binning, benchmarked against the XGBoost model. Class imbalance handled with SMOTE and cost-sensitive thresholds, which notably improved minority-class recall; PSI-based drift monitoring with MLflow versioning.

  • Logistic Regression
  • WOE/IV
  • SMOTE
  • MLflow

03

PySpark · AWS Glue · CDC

Enterprise data platform

Design, build and own 8+ ETL/ELT pipelines (PySpark, Lambda, Glue) processing 4B+ rows across 5+ source systems, using CDC and a medallion architecture. Shipped the Data Warehouse 3 months ahead of deadline.

  • PySpark
  • AWS Glue
  • CDC
  • Medallion

04

NL-to-query · LLM summaries

Natural-language analytics

Added a natural-language-to-query layer to the Lending Analytics Platform (5,000+ employees, live production data) that turns plain-English questions into LLM-generated trend summaries.

  • NL-to-SQL
  • LLM
  • Analytics

05

QLoRA · benchmarking

Domain LLM fine-tune

Fine-tuned an open-source LLM with QLoRA for a lending-domain task and benchmarked it against prompted baselines on accuracy, latency and cost.

  • QLoRA
  • Evaluation
  • Transformers

06

PostgreSQL · indexing · query plans

Query performance

Tuned PostgreSQL queries behind the MIS dashboard through indexing and query-plan work, cutting report latency from 2 hours to 15 minutes.

  • PostgreSQL
  • Performance
  • BI

Toolkit

ML & statistics

Classification, clustering, WOE/IV scorecards, SHAP explainability, hypothesis testing, survival analysis, imbalanced classification (SMOTE), feature engineering

GenAI & applied analytics

RAG pipelines, fine-tuning, LLM-generated reporting, natural-language-to-query analytics

Tools & libraries

Python, XGBoost, scikit-learn, PyTorch, Transformers, statsmodels, LangChain/LangGraph, AWS Bedrock, embeddings, Qdrant

Data analysis & BI

SQL, Spark SQL, PostgreSQL, Power BI, query profiling and tuning, exploratory data analysis, dashboarding

Data engineering

PySpark, AWS Glue, CDC, medallion architecture, ETL/ELT design and ownership, Kafka, Airflow, data warehousing and modelling

Engineering & cloud

FastAPI, REST APIs, Docker, Redis, AWS Lambda, Redshift, EMR, Step Functions, RDS

Experience

Oct 2025 — present

Data and AI Engineer

AYE Finance

Build the ML models and scorecards behind credit and collections decisions, and own the AWS data pipelines that feed them. Lead 6–7 engineers across data, ML and analytics initiatives through sprint planning, code review and stakeholder delivery.

Dec 2024 — Sep 2025

IT Intern — Data Warehousing

AYE Finance

Led end-to-end design of a production data warehouse and ETL pipelines handling 10M+ records across LOS, LMS and CMS using SQL, PySpark and Python. Mapped complete data lineage, presented the architecture to the CTO and delivered an RBI compliance app 3× faster than planned with a 10-member team. Awarded “Cha Gaye — February” for delivery impact.

Aug 2024 — Dec 2024

AI Intern

Accodyn Technologies

Built an ML-based risk detection model reaching 93% accuracy for a lending application, and contributed full-stack ML features and UI/UX.

Jan 2024 — Jun 2024

Technical Intern

RK Agrotech

Built an inventory and lead-management dashboard (HTML/CSS/JS + FastAPI) backed by PostgreSQL, with automated data ingestion from Google Forms and Sheets.

Education

2026 — 2028

M.Sc. in Data Science & AI

BITS Pilani (Digital)

In progress.

2021 — 2025

B.Tech. in Computer Science & Engineering

MIT-WPU, Pune

Recognition

  • 2nd Runner-up, AWS Agentic AI Hackathon
  • 6+ internal awards for delivery excellence and impact
  • HackerRank SQL 5★

Selected projects

all 10 projects →

Collections Prioritization model

XGBoost · WOE/IV · SMOTE · MLflow · LLM

An XGBoost model that ranks overdue loan accounts by recovery likelihood, with an interpretable scorecard benchmark and LLM-written segment reports.

view on engineering hub →

Credit Scorecard & Survival-Causal Study

Python · scikit-learn · statsmodels · lifelines · Credit Risk

An interpretable WOE/IV credit scorecard benchmarked against gradient boosting, plus survival and causal analysis of time-to-default on a public lending dataset.

view on engineering hub →

LoanRoute Pipeline

XGBoost · SHAP · FastAPI · ML

Prioritises loan applicants with an XGBoost model and explains each ranking with SHAP.

view on engineering hub →

DataFlow

ETL · Python · DuckDB · Spark · YAML

A metadata-driven ETL framework that compiles YAML pipeline definitions into DuckDB, Polars or Spark jobs.

view on engineering hub →

Operations dashboard

React · FastAPI · PostgreSQL · Analytics · RBAC

A live KPI dashboard for branch, zone and state operations at AYE Finance: React on the front, FastAPI and PostgreSQL behind it.

view on engineering hub →

Credit Helper Agent & policy assistant

LangGraph · RAG · AWS Bedrock · Qdrant · FastAPI · Electron

Multi-agent RAG for credit and policy questions at AYE Finance, with guardrails, private document chat, and desktop and WhatsApp clients.

view on engineering hub →

Technical articles

4 articles