Backend systems at scale
Java, Spring Boot, event pipelines, observability, RAG, and agentic AI systems.
- Focus
- Backend · AI / ML systems
- Core stack
- Java · Spring Boot · Python · PostgreSQL
- AI stack
- LangChain · LangGraph · RAG · Vector search
- Based in
- Gurugram, India
Education
Master of Computer Applications
Hindustan University, Chennai · Sep 2022 — Sep 2024 · CGPA 8.4
Active trajectory
Airtel Digital
Software Engineering Intern
Jan 2025 — Jan 2026
Gurugram, India
- N16 services / 25+ endpointsBuilt and maintained backend microservices and REST endpoints in Java, Spring Boot and PostgreSQL, supporting enterprise analytics workflows used by 4 internal teams.
- N22 hours → under 5 minutesDeveloped AI-powered backend workflows with LLM APIs, LangChain and a RAG retrieval layer, cutting enterprise document lookup per request.
- N312M events / day at 99.5% uptimeDesigned distributed Apache Spark and Kafka pipelines ingesting telemetry, reducing batch processing time by 35%.
- N490% monitoring coverage, −40% MTTDInstrumented 8 services with Prometheus, Grafana, Loki, Tempo and OTLP tracing, cutting mean time to diagnose incidents.
- N545 min → under 10 min releasesAutomated containerised deployments with Docker, Kubernetes and CI/CD, cutting failed deployments by 60%.
- N6850 ms → 300 ms p95Optimised SQL queries and service APIs under production load.
Problem
A school deciding on outdoor practice has one current AQI number to go on: nothing station-level for the days ahead, and public feeds that go silent without saying so.
Approach
Hourly ingestion from three sources into TimescaleDB, orchestrated by nine Airflow DAGs. Each of the next five days is graded on the official CPCB scale and served by a FastAPI API to a React dashboard, with a watchdog, tested backups and a nightly evaluation behind it.
Trade-off
Served a simple statistical rule instead of the gradient-boosted models I had already built. In block-holdout backtests no model beat "the next days look like the last 24 hours", so I kept the option that was as accurate, better calibrated and unable to learn a sensor fault.
What broke / what I'd change
A sensor reading of 2.9 million µg/m³ became a training label and the API served a forecast of 8,243. I now look at the extremes of every new data source before anything is computed from it.
Result
Live for about 80 stations. In backtests the grade is exactly right on about 6 in 10 days for tomorrow, and a missing forecast is never shown as "go".
Python / Apache Airflow / PostgreSQL + TimescaleDB / FastAPI / Docker / TypeScript / React
Open project page →Live dashboard ↗Problem
Job applications are a long chain of small, stateful tasks that break the moment a single-prompt assistant loses context.
Approach
Graph-based multi-agent orchestration with tool calling and structured memory, persisting state between agent steps so a run can be resumed and inspected.
Trade-off
Chose an explicit state graph over a single autonomous agent loop: more wiring and more code per capability, in exchange for runs that can be replayed and debugged step by step.
What broke / what I'd change
Long runs drifted when a tool returned an unexpected shape. Next pass: schema-validate every tool result at the graph edge and fail the node instead of letting the model improvise around it.
Result
End-to-end task execution instead of one-shot suggestions.
Python / LangGraph / LLM APIs / Tool Calling
Open project page →Problem
Answers lived inside thousands of internal documents, and keyword search kept returning the wrong page.
Approach
Chunking, embedding, retrieval and grounded answer generation with source citations, exposed as a REST API other services can call.
Trade-off
Kept retrieval to dense vector search with citations rather than a reranking stack: slightly weaker on ambiguous phrasing, but predictable latency and a much smaller surface to operate.
What broke / what I'd change
Fixed-size chunking split tables down the middle and produced confident wrong answers. Structure-aware chunking is the first thing I would rebuild.
Result
Document lookup dropped from around two hours to under five minutes.
Python / LangChain / Vector store / REST API
Open project page →Problem
Analysts waited on engineers for every ad-hoc query.
Approach
Schema-aware SQL generation from plain English inside Apache Superset, with validation before execution so nothing unsafe reaches the warehouse.
Trade-off
Validation rejects anything it cannot prove safe, so some legitimate queries get refused. Blocking a valid question is cheaper than letting a generated statement touch production data.
What broke / what I'd change
Generation quality tracked schema documentation more than model choice. I would invest in column descriptions and query examples before touching prompts again.
Result
Self-serve querying without writing SQL.
Python / LLM APIs / Apache Superset / SQL
Open project page →Problem
High-volume device telemetry arriving faster than batch jobs could absorb it.
Approach
Fault-tolerant streaming from Kafka into Spark for transformation and aggregation, with checkpointing and recovery on failure.
Trade-off
Micro-batching over per-event streaming: seconds of added latency, in return for throughput and recovery behaviour that survives a node dying mid-window.
What broke / what I'd change
Checkpoint state grew until restarts got slow. Retention and compaction belong in the design from day one, not in a later fix.
Result
~12M events a day sustained at 99.5% uptime.
Apache Spark / Kafka / Java / Docker
Open project page →Node 01 / Languages
Java · Python · SQL · JavaScript · Bash
Node 02 / Backend & Data
Spring Boot · REST APIs · Microservices · Apache Spark · Apache Kafka
Node 03 / AI & LLMs
RAG pipelines · LangChain · LangGraph · Agentic workflows · Tool calling · MCP · Embeddings & vector search · Prompt engineering
Node 04 / Databases
PostgreSQL · MongoDB · OracleDB
Node 05 / Cloud & Ops
AWS · Docker · Kubernetes · Linux · CI/CD · Prometheus · Grafana · Loki · Tempo · OpenTelemetry
- 01Emailhimanshubaliyan4000@gmail.com
- 02Phone+91 70068 02968
- 03LinkedInlinkedin.com/in/himanshu-baliyan
- 04RésuméDownload PDF · September 2026