DEVOPS · CLOUD · PLATFORM ENGINEERING

Amit Gadhia

9+ years designing and operating cloud-native infrastructure, Kubernetes platforms, and CI/CD pipelines across AWS, GCP, and Azure — with a specialisation in MLOps and GPU infrastructure for teams shipping ML to production.

9+
Years Experience
50+
Projects Delivered
7+
Cloud Platforms
amit@platform: ~
$ terraform apply -auto-approve
Apply complete!  42 added, 0 destroyed
 
$ kubectl get nodes
prod-cp-1    Ready  control-plane  v1.30
prod-worker-3  Ready   <none>     v1.30
gpu-worker-2   Ready   gpu        v1.30
 
$ nvidia-smi --query-gpu=utilization.gpu
GPU 0 · H100 80GB0%
GPU 1 · A100 40GB0%
 
$ echo "infra as code · models in prod"
infra as code · models in prod
$

Who I Am

Amit Gadhia — DevOps, Cloud & Platform Engineering Architect $ whoami → amit

I'm a hands-on DevOps, Cloud & Platform Engineering Architect with over 9 years of experience turning complex business challenges into scalable, resilient technology platforms — from cloud infrastructure and Kubernetes to CI/CD and automation.

My career spans high-growth startups and large enterprise environments, where I've designed cloud-native architectures, built CI/CD and GitOps pipelines, hardened production systems, and modernised legacy applications into cloud-native microservices across AWS, GCP, and Azure.

In recent years I've extended that platform foundation into MLOps, GPU infrastructure, and AI agents — orchestrating GPU workloads on Kubernetes, building the pipelines that let teams ship models as confidently as they ship code, and deploying, scaling, and securing agentic systems in production.

DevOps Cloud Architecture Platform Engineering Kubernetes Infrastructure as Code CI/CD & GitOps MLOps GPU Infrastructure AI Agents
Email
amitgadhia65@gmail.com
Location
Jammu, India
Experience
9+ Years
Status
Available for Freelance

What I Do Best

Core capabilities built over 9 years across cloud infrastructure, Kubernetes, and DevOps automation — extending into MLOps and GPU compute.

Kubernetes & Cloud

Designing multi-region, multi-cloud Kubernetes platforms on AWS, GCP, and Azure — networking, security hardening, high availability, and cost optimisation.

EKSGKEAKSHelm

CI/CD, GitOps & Automation

Automated delivery pipelines from commit to production — GitOps workflows, progressive delivery, release automation, and self-service platforms that let teams ship safely and often.

GitLab CIArgoCDJenkinsGitHub Actions

Infrastructure as Code

Everything versioned and reproducible — Terraform modules, configuration management, and automated environments that spin up identical stacks in minutes.

TerraformAnsibleChefPacker

Observability & Reliability

Full-stack visibility across applications and infrastructure — metrics, logging, tracing, alerting, and SLO-driven operations that keep production healthy.

PrometheusGrafanaEFKOpenTelemetry

GPU & Accelerated Compute

Provisioning, scheduling, and optimising GPU workloads on Kubernetes — node pools, GPU sharing, autoscaling, and cost control for training and inference fleets.

CUDAGPU OperatorDCGMMIG

MLOps & LLM Serving

End-to-end ML platforms and production model serving — experiment tracking, model registries, automated retraining, and high-throughput LLM inference with vLLM and Triton.

KubeflowMLflowvLLMTriton

Where I've Worked

A track record of building and scaling infrastructure across startups and enterprise projects.

DevOps Engineer 2018 – 2019
Indivar Software Solutions

Designed and maintained multi-region Kubernetes clusters on DigitalOcean. Implemented full CI/CD pipelines using GitLab CI for containerised application deployments. Established production observability with the EFK stack and provided cross-team support through testing and production phases.

Kubernetes GitLab CI Docker AWS Lambda EFK Stack
Systems Engineer 2017 – 2018
Dbaux Pvt Ltd

Automated patch management and distributed server operations using Bash and Python scripting. Managed infrastructure spanning 10+ production servers and 200–300 client machines. Implemented configuration management with Chef and deployed EFK stack monitoring across the fleet.

Python Bash Chef EFK Stack Linux Automation

How I Can Help

End-to-end engineering services for teams building and operating cloud infrastructure at scale — and shipping ML to production.

01

DevOps & Cloud Architecture

Multi-cloud strategy and migration across AWS, GCP, and Azure — cloud-native architecture, legacy application modernisation into microservices, security hardening, high availability, and cost optimisation.

02

Platform Engineering & CI/CD

Kubernetes platforms and delivery pipelines that scale with your team — GitOps workflows, Infrastructure as Code, self-service developer platforms, progressive delivery, and end-to-end observability.

03

MLOps Platform Design

End-to-end ML infrastructure — training pipelines, experiment tracking, feature stores, model registries, and automated promotion from experimentation to production with full lineage and reproducibility.

04

GPU & LLM Inference

GPU compute platforms and production model serving — driver and operator management, GPU sharing and MIG partitioning, vLLM and Triton deployments, throughput and latency tuning, and cost optimisation for training and inference fleets.

Agents in Production

Taking LLM and multi-agent systems from prototype to production — deployed reliably, scaled under real load, and secured with least-privilege guardrails.

Deployment & Orchestration

Shipping agentic systems to production — containerised runtimes, tool and function-calling infrastructure, MCP servers, multi-agent orchestration, state and memory backends, and CI/CD for prompts, tools, and models.

LangGraphMCPKubernetesRay Serve

Scaling & Reliability

Holding up under real traffic — concurrency and queueing for long-running agent runs, event-driven autoscaling, token and GPU cost control, caching and batching, timeouts, retries, graceful degradation, and step-level tracing.

KEDAAutoscalingRedisOpenTelemetry

Security & Governance

Least-privilege by design — sandboxed tool execution, scoped credentials and secrets, network egress controls, prompt-injection and jailbreak defenses, human-in-the-loop approvals, rate limiting, PII redaction, and full audit logging.

SandboxingGuardrailsRBACAudit Logs

Tools I Work With

The platforms and tooling I use daily to build and operate cloud infrastructure — and the ML stack on top of it.

Amazon Web Services
Google Cloud
Microsoft Azure
Kubernetes
Docker
Terraform
GitLab CI
Jenkins
ELK Stack
Prometheus
Grafana
Python
Bash / Shell
Git
REST APIs
NVIDIA CUDA
Triton Server
vLLM
Ray
PyTorch
KServe
Kubeflow
MLflow
Apache Airflow
LangGraph
MCP

What Clients Say

"Amit is a great worker who delivered outstanding results. His technical depth and reliability are remarkable. We 100% recommend him to anyone looking for exceptional infrastructure work."

Zach Rattner
Zach Rattner
Client

"Amit is a terrific technical resource, a great individual contributor, and a genuinely kind person. You cannot go wrong by working with him on any project."

Mike Pollard
Mike Pollard
Client

"Very good work and always available. Delivered excellent output in record time with clear communication throughout the entire engagement."

Vijayababu Emani
Vijayababu Emani
Client

Get In Touch

Have cloud infrastructure to build, a platform to modernise, or ML to ship to production? Let's talk about how I can help.

Email
amitgadhia65@gmail.com
Location
Jammu, India
Availability
Open to freelance & consulting