AI that ships.
RAG pipelines, AI agents, embeddings, and local inference. Generative AI integrated into real production workflows.
I’m Shubhankar Kumar Singh.
I build AI-native products, real-time experiences, and
distributed systems that hold up beyond the demo.
Full-stack engineering experience
Concurrent users · AI chat platform
Notification events processed daily
Production uptime achieved
RAG pipelines, AI agents, embeddings, and local inference. Generative AI integrated into real production workflows.
Event-driven services, asynchronous processing, distributed rate limiting, and high-concurrency real-time platforms.
Secure APIs, responsive interfaces, cloud deployments, and observability. Built with reliability in mind from the start.
Focused tools for difficult engineering problems. Built for concurrency, practical AI, and reuse.
A distributed, atomic token-bucket rate limiter for Node.js. Redis Lua scripts keep enforcement consistent across horizontally scaled instances—even under heavy concurrency.
An OpenAI-compatible embeddings server powered by Node.js and ONNX Runtime. Run transformer models locally on CPUs, without external inference APIs or network round trips.
From the interface to the inference pipeline. Choose a category, then select a technology to explore.
JavaScript, TypeScript, Python, SQL, Java, React.js, Redux Toolkit, Tailwind CSS, responsive UI, state management, frontend–backend integration.
Node.js, Express.js, FastAPI, RESTful APIs, GraphQL, WebSockets, SSE, microservices, distributed systems, asynchronous processing, high availability.
Generative AI, LLM integration, AI agents, RAG, semantic search, embeddings, LangChain, OpenAI, Azure OpenAI, Pinecone, MongoDB Atlas Vector Search, ONNX Runtime.
PostgreSQL, MongoDB, MySQL, Redis, RabbitMQ, BullMQ, data modeling, schema design, query optimization, retries, delivery guarantees, distributed rate limiting.
AWS EC2, S3, Lambda, SES, Docker, CI/CD, Firebase, Cloudflare Workers, Grafana, Loki, Sentry, Git, production monitoring, incident troubleshooting, RCA.
RBAC, authentication, authorization, input validation, secure API design, reCAPTCHA Enterprise, concurrency tuning, latency optimization, throughput optimization, SDK development.
Selected outcomes from production engineering. Different systems. The same focus on measurable improvement.
Deepgram streaming speech-to-text, asynchronous processing, and Node.js concurrency optimization for a real-time voice pipeline.
A WebSocket and LangChain AI chat platform supporting 10K+ concurrent users through high-concurrency distributed processing.
Centralized logging, metrics, and production monitoring with Grafana and Loki to support troubleshooting and root-cause analysis.
From cloud infrastructure to enterprise AI. Five years of building, optimizing, debugging, and shipping across the entire application stack.
Connect on LinkedIn ↗
Netaji Subhash Engineering College
Kolkata, West Bengal · July 2016 — July 2020
Experience across remote US teams and Indian technology companies.
Building an AI product, scaling a platform, or looking for an engineer who owns the outcome? Let’s start a conversation.
shubhankars361@gmail.com