9+ years designing and operating cloud-native infrastructure, Kubernetes platforms, and CI/CD pipelines across AWS, GCP, and Azure — with a specialisation in MLOps and GPU infrastructure for teams shipping ML to production.
$ whoami → amit
I'm a hands-on DevOps, Cloud & Platform Engineering Architect with over 9 years of experience turning complex business challenges into scalable, resilient technology platforms — from cloud infrastructure and Kubernetes to CI/CD and automation.
My career spans high-growth startups and large enterprise environments, where I've designed cloud-native architectures, built CI/CD and GitOps pipelines, hardened production systems, and modernised legacy applications into cloud-native microservices across AWS, GCP, and Azure.
In recent years I've extended that platform foundation into MLOps, GPU infrastructure, and AI agents — orchestrating GPU workloads on Kubernetes, building the pipelines that let teams ship models as confidently as they ship code, and deploying, scaling, and securing agentic systems in production.
Core capabilities built over 9 years across cloud infrastructure, Kubernetes, and DevOps automation — extending into MLOps and GPU compute.
Designing multi-region, multi-cloud Kubernetes platforms on AWS, GCP, and Azure — networking, security hardening, high availability, and cost optimisation.
Automated delivery pipelines from commit to production — GitOps workflows, progressive delivery, release automation, and self-service platforms that let teams ship safely and often.
Everything versioned and reproducible — Terraform modules, configuration management, and automated environments that spin up identical stacks in minutes.
Full-stack visibility across applications and infrastructure — metrics, logging, tracing, alerting, and SLO-driven operations that keep production healthy.
Provisioning, scheduling, and optimising GPU workloads on Kubernetes — node pools, GPU sharing, autoscaling, and cost control for training and inference fleets.
End-to-end ML platforms and production model serving — experiment tracking, model registries, automated retraining, and high-throughput LLM inference with vLLM and Triton.
A track record of building and scaling infrastructure across startups and enterprise projects.
Designed and maintained multi-region Kubernetes clusters on DigitalOcean. Implemented full CI/CD pipelines using GitLab CI for containerised application deployments. Established production observability with the EFK stack and provided cross-team support through testing and production phases.
Automated patch management and distributed server operations using Bash and Python scripting. Managed infrastructure spanning 10+ production servers and 200–300 client machines. Implemented configuration management with Chef and deployed EFK stack monitoring across the fleet.
End-to-end engineering services for teams building and operating cloud infrastructure at scale — and shipping ML to production.
Multi-cloud strategy and migration across AWS, GCP, and Azure — cloud-native architecture, legacy application modernisation into microservices, security hardening, high availability, and cost optimisation.
Kubernetes platforms and delivery pipelines that scale with your team — GitOps workflows, Infrastructure as Code, self-service developer platforms, progressive delivery, and end-to-end observability.
End-to-end ML infrastructure — training pipelines, experiment tracking, feature stores, model registries, and automated promotion from experimentation to production with full lineage and reproducibility.
GPU compute platforms and production model serving — driver and operator management, GPU sharing and MIG partitioning, vLLM and Triton deployments, throughput and latency tuning, and cost optimisation for training and inference fleets.
Taking LLM and multi-agent systems from prototype to production — deployed reliably, scaled under real load, and secured with least-privilege guardrails.
Shipping agentic systems to production — containerised runtimes, tool and function-calling infrastructure, MCP servers, multi-agent orchestration, state and memory backends, and CI/CD for prompts, tools, and models.
Holding up under real traffic — concurrency and queueing for long-running agent runs, event-driven autoscaling, token and GPU cost control, caching and batching, timeouts, retries, graceful degradation, and step-level tracing.
Least-privilege by design — sandboxed tool execution, scoped credentials and secrets, network egress controls, prompt-injection and jailbreak defenses, human-in-the-loop approvals, rate limiting, PII redaction, and full audit logging.
The platforms and tooling I use daily to build and operate cloud infrastructure — and the ML stack on top of it.
"Amit is a great worker who delivered outstanding results. His technical depth and reliability are remarkable. We 100% recommend him to anyone looking for exceptional infrastructure work."
"Amit is a terrific technical resource, a great individual contributor, and a genuinely kind person. You cannot go wrong by working with him on any project."
"Very good work and always available. Delivered excellent output in record time with clear communication throughout the entire engagement."
Have cloud infrastructure to build, a platform to modernise, or ML to ship to production? Let's talk about how I can help.