No chapter in view.
01Orbital designation / software engineer
HimanshuBaliyan
I build production backend services in Java and Spring Boot, and ship LLM-powered systems: RAG pipelines, agentic workflows and the data plumbing underneath them.
Download résuméView projectsUpdated September 2026
System live28.6139° N / 77.2090° E
Current focusActive

Backend systems at scale

Java, Spring Boot, event pipelines, observability, RAG, and agentic AI systems.

02Profile / system architecture
Systemsthat holdup.
Designing resilient services and intelligent workflows where reliability matters as much as invention.
Software engineer with a year of production experience across backend services and AI systems. Most of my work sits where the two meet: REST services and data pipelines on one side, retrieval and agent orchestration on the other.
I care about the unglamorous parts — p95 latency, pipeline uptime, tracing coverage, deploy time — because those are what decide whether a feature survives contact with real traffic.
Focus
Backend · AI / ML systems
Core stack
Java · Spring Boot · Python · PostgreSQL
AI stack
LangChain · LangGraph · RAG · Vector search
Based in
Gurugram, India

Education

Master of Computer Applications

Hindustan University, Chennai · Sep 2022 — Sep 2024 · CGPA 8.4

03Distributed operations
Productionunder load.
12,000,000 events / 24h

Active trajectory

Airtel Digital

Software Engineering Intern

Jan 2025 — Jan 2026
Gurugram, India

  1. N16 services / 25+ endpointsBuilt and maintained backend microservices and REST endpoints in Java, Spring Boot and PostgreSQL, supporting enterprise analytics workflows used by 4 internal teams.
  2. N22 hours → under 5 minutesDeveloped AI-powered backend workflows with LLM APIs, LangChain and a RAG retrieval layer, cutting enterprise document lookup per request.
  3. N312M events / day at 99.5% uptimeDesigned distributed Apache Spark and Kafka pipelines ingesting telemetry, reducing batch processing time by 35%.
  4. N490% monitoring coverage, −40% MTTDInstrumented 8 services with Prometheus, Grafana, Loki, Tempo and OTLP tracing, cutting mean time to diagnose incidents.
  5. N545 min → under 10 min releasesAutomated containerised deployments with Docker, Kubernetes and CI/CD, cutting failed deployments by 60%.
  6. N6850 ms → 300 ms p95Optimised SQL queries and service APIs under production load.
04Selected orbits
Systemsin motion.
Each shipped system occupies its own orbit. Every entry carries the problem, the decision made against something else, and what it cost.

Problem

A school deciding on outdoor practice has one current AQI number to go on: nothing station-level for the days ahead, and public feeds that go silent without saying so.

Approach

Hourly ingestion from three sources into TimescaleDB, orchestrated by nine Airflow DAGs. Each of the next five days is graded on the official CPCB scale and served by a FastAPI API to a React dashboard, with a watchdog, tested backups and a nightly evaluation behind it.

Trade-off

Served a simple statistical rule instead of the gradient-boosted models I had already built. In block-holdout backtests no model beat "the next days look like the last 24 hours", so I kept the option that was as accurate, better calibrated and unable to learn a sensor fault.

What broke / what I'd change

A sensor reading of 2.9 million µg/m³ became a training label and the API served a forecast of 8,243. I now look at the extremes of every new data source before anything is computed from it.

Result

Live for about 80 stations. In backtests the grade is exactly right on about 6 in 10 days for tomorrow, and a missing forecast is never shown as "go".

Python / Apache Airflow / PostgreSQL + TimescaleDB / FastAPI / Docker / TypeScript / React

Open project page →Live dashboard ↗

Problem

Job applications are a long chain of small, stateful tasks that break the moment a single-prompt assistant loses context.

Approach

Graph-based multi-agent orchestration with tool calling and structured memory, persisting state between agent steps so a run can be resumed and inspected.

Trade-off

Chose an explicit state graph over a single autonomous agent loop: more wiring and more code per capability, in exchange for runs that can be replayed and debugged step by step.

What broke / what I'd change

Long runs drifted when a tool returned an unexpected shape. Next pass: schema-validate every tool result at the graph edge and fail the node instead of letting the model improvise around it.

Result

End-to-end task execution instead of one-shot suggestions.

Python / LangGraph / LLM APIs / Tool Calling

Open project page →

Problem

Answers lived inside thousands of internal documents, and keyword search kept returning the wrong page.

Approach

Chunking, embedding, retrieval and grounded answer generation with source citations, exposed as a REST API other services can call.

Trade-off

Kept retrieval to dense vector search with citations rather than a reranking stack: slightly weaker on ambiguous phrasing, but predictable latency and a much smaller surface to operate.

What broke / what I'd change

Fixed-size chunking split tables down the middle and produced confident wrong answers. Structure-aware chunking is the first thing I would rebuild.

Result

Document lookup dropped from around two hours to under five minutes.

Python / LangChain / Vector store / REST API

Open project page →

Problem

Analysts waited on engineers for every ad-hoc query.

Approach

Schema-aware SQL generation from plain English inside Apache Superset, with validation before execution so nothing unsafe reaches the warehouse.

Trade-off

Validation rejects anything it cannot prove safe, so some legitimate queries get refused. Blocking a valid question is cheaper than letting a generated statement touch production data.

What broke / what I'd change

Generation quality tracked schema documentation more than model choice. I would invest in column descriptions and query examples before touching prompts again.

Result

Self-serve querying without writing SQL.

Python / LLM APIs / Apache Superset / SQL

Open project page →

Problem

High-volume device telemetry arriving faster than batch jobs could absorb it.

Approach

Fault-tolerant streaming from Kafka into Spark for transformation and aggregation, with checkpointing and recovery on failure.

Trade-off

Micro-batching over per-event streaming: seconds of added latency, in return for throughput and recovery behaviour that survives a node dying mid-window.

What broke / what I'd change

Checkpoint state grew until restarts got slow. Retention and compaction belong in the design from day one, not in a later fix.

Result

~12M events a day sustained at 99.5% uptime.

Apache Spark / Kafka / Java / Docker

Open project page →
05Dependency constellation
Built forthe wholesystem.

Node 01 / Languages

Java · Python · SQL · JavaScript · Bash

Node 02 / Backend & Data

Spring Boot · REST APIs · Microservices · Apache Spark · Apache Kafka

Node 03 / AI & LLMs

RAG pipelines · LangChain · LangGraph · Agentic workflows · Tool calling · MCP · Embeddings & vector search · Prompt engineering

Node 04 / Databases

PostgreSQL · MongoDB · OracleDB

Node 05 / Cloud & Ops

AWS · Docker · Kubernetes · Linux · CI/CD · Prometheus · Grafana · Loki · Tempo · OpenTelemetry

End of telemetry — open communication vector below
06 / Open communication vector — Gurugram, India
Himanshu
Let’s build what lasts.
  1. 01Emailhimanshubaliyan4000@gmail.com
  2. 02Phone+91 70068 02968
  3. 03LinkedInlinkedin.com/in/himanshu-baliyan
  4. 04RésuméDownload PDF · September 2026