Raj Aryan · Backend & distributed systems

Systems that keep their shape under real-world traffic.

I design and operate high-throughput backend systems where milliseconds, failure modes and infrastructure cost directly affect product outcomes.

Open to senior backend, platform and selected architecture-consulting conversations.

Selected operating scale

Evidence, not decoration.

A few outcomes from production systems where performance and reliability are product concerns.

4.3M+

requests per minute operated across high-scale backend systems.

7M+

Kafka events processed daily in an event-driven pipeline.

500 → 240 ms

location-search latency after a service redesign.

93%

reduction in database-query volume through caching improvements.

≈3%

order-rate improvement from real-time promotion personalisation.

Engineering focus

Start with the problem, then choose the technology.

Scale & performance

  • High-throughput APIs
  • Go profiling & allocation analysis
  • Latency and capacity analysis
  • Caching strategy

Reliability

  • Failure-mode analysis
  • Retries, timeouts & backoff
  • Graceful degradation
  • Observability & incident analysis

Data & event systems

  • Kafka pipelines
  • Redis and cache boundaries
  • DynamoDB and SQL systems
  • Stream processing

Search & discovery

  • Solr, Lucene & Elasticsearch
  • Geo-sharding
  • Indexing pipelines
  • Fallback search paths

Infrastructure

  • AWS and containers
  • Service communication
  • Deployment safety
  • Cloud-cost optimisation

Product systems

  • Promotions and rewards
  • Personalisation
  • Search, homepage & location
  • Business-aware trade-offs

Selected engagements

A focused second pair of eyes on a system that matters.

Explore services
01

Backend architecture risk review

A structured walkthrough of the decisions most likely to constrain scale or reliability.

  • Scalability-risk assessment
  • Reliability gaps
  • Prioritised recommendations
  • Written action plan
02

Production reliability audit

Trace how dependency failures become customer impact—and identify controls that break the chain.

  • Retry and timeout review
  • Failure-mode map
  • Observability gaps
  • Degraded-mode strategy
03

Performance & cost review

Find the hotspots worth measuring before a performance or cloud-cost problem becomes structural.

  • CPU and memory hotspots
  • Query and cache efficiency
  • Throughput bottlenecks
  • Measurement plan

Writing

Notes from building systems that must keep working.

Read the writing
MEDIUM · PRODUCTION RELIABILITY · 4 MIN READ

Retry Storm: How A Single User Crashed 30 ECS Tasks

A production debugging story about infrastructure-level retry amplification and 503 ambiguity.

Read on Medium
MEDIUM · SEARCH INFRASTRUCTURE · 8 MIN READ

When Apache Solr’s Replication Handler Went Rogue

A production debugging story about replication timing, stale commit points and leader changes.

Read on Medium
DRAFT · CACHING · 5 MIN READ

Cache boundaries are an API design decision

How to reason about cache ownership, invalidation and degraded reads.

Read draft

Labs, earlier builds & open source

Small experiments with a life beyond the day job.

A small set of projects from Raj’s public GitHub work—useful context alongside the production systems above.

PUBLIC TOOLING · 40 STARS · 18 FORKS

Career Copilot

An AI-powered job-search pipeline for coding agents, built as a public template and workflow.

View repository
FLUTTER · 66 STARS · 25 FORKS

Vintage Pacman in Flutter

A mobile take on Pacman that was also featured by Flutter Awesome.

View repository
PYTHON · 25 STARS · 7 FORKS

Hindi Text-to-Speech

A phoneme-database approach to Hindi text-to-speech, shared as a public Python project.

View repository

GitHub popularity. Snapshot recorded 29 July 2026. Personal GitHub · Work GitHub

Raj Aryan in a yellow kurta

About

Calm systems are built by treating constraints as first-class inputs.

“The interesting work starts where a clean happy path meets real traffic, imperfect dependencies and an operating budget.”

Raj is a backend engineer focused on distributed systems that serve product teams and customers under load. The work spans promotions, rewards, subscriptions, search, homepage, dish and location platforms.

A practical next step

Have a backend problem that is too important to guess at?

Share the system, scale and outcome you are working toward. I’ll respond when the problem is a strong fit.