Ravikumar M
Cloud Engineering Specialist
10 years of experience architecting scalable cloud-native platforms, building AI-powered operational automation, and ensuring high availability across multi-cloud environments.
About Me
DevOps, Cloud & SRE Engineer with 10 years of experience architecting scalable cloud-native platforms, building AI-powered operational automation, and ensuring high availability across multi-cloud environments.
Currently driving AI adoption across the software lifecycle — implementing AI-assisted development, testing, and automation. Leading cloud-native software engineering with emphasis on scalability, resilience, security, and fault tolerance. Ensuring platform reliability through observability, incident management, and continuous improvement.
Deep expertise in Kubernetes, Infrastructure as Code, CI/CD, and Linux administration, combined with hands-on experience designing and deploying RAG-based AI systems, Generative AI tooling, and intelligent automation using OpenSearch and LLMs. Strong DevSecOps practitioner — embedding security into CI/CD pipelines, container scanning, runtime protection, and compliance hardening.
Proven track record of integrating AI into DevOps workflows — including automated incident remediation, AI-driven data flow generation, and intelligent deployment validation. Building enterprise-scale platforms handling 10M+ daily transactions at BT Group.
Current Role
Cloud Engineering Specialist
Focus Areas
DevOps · Cloud · SRE · DevSecOps · AIOps · MLOps
Location
Bengaluru, Karnataka
Experience
10+ Years
Technical Skills
AI/ML & Gen AI
- RAG Pipelines
- LLM Fine-Tuning
- AI Agents & Agentic AI
- Amazon Bedrock
- OpenSearch Vector DB
- Prompt Engineering
- LangChain, LangGraph, CrewAI
- Ollama, Gemma 3, Nomic Embed
MLOps
- SageMaker Pipelines & Studio
- Model Registry & Serving
- DVC & MLflow
- KServe
- Model Monitoring
- Experiment Tracking
Cloud Platforms
- AWS (EKS, Lambda, S3, ECR, IAM, SQS, VPC, SES, Step Functions, Bedrock)
- Azure (AKS)
- Multi-Cloud Architecture
- Cost Optimisation & Capacity Planning
Containers & Orchestration
- Kubernetes (EKS, AKS, GKE)
- Docker & Container Builds
- Helm Charts
- ArgoCD & GitOps
- Istio Service Mesh
- Cluster Upgrades & Patching
IaC & CI/CD
- Terraform & Pulumi
- GitLab CI/CD & Jenkins
- GitOps Workflows
- Blue/Green & Canary Deployments
- Release Governance & Rollbacks
SRE & Observability
- Dynatrace (Full-Stack APM)
- Prometheus & Grafana
- ELK Stack & Distributed Tracing
- Incident Management & RCA
- SLA/SLO/SLI Management
- Resilience Engineering
- Capacity Planning
DevSecOps & Security
- Container Security (StackRox)
- Image Vulnerability Scanning
- CIS Benchmark & STIG Compliance
- Security Patching & CVE Remediation
- Runtime Threat Prevention
- Secrets Management
Linux & Infrastructure
- Linux (RHEL, Ubuntu, CentOS, Debian)
- HA Load Balancing (Nginx, KeepAlived)
- DNS, DHCP, NTP, SMTP, Proxy
- Virtualisation (KVM, XEN, OpenVZ)
- Disaster Recovery & Backup
- 99.99% Uptime SLA
Scripting & Data Platforms
- Python & Bash
- Apache NiFi (Admin & Development)
- Apache Kafka
- AWS SQS & Event-Driven Architecture
- Puppet, Ansible, Cobbler
Work Experience
Cloud Engineering Specialist
BT Group
Dec 2023 – Present
Bengaluru, Karnataka
Driving AI adoption across the software lifecycle through AI-assisted development, testing, and automation across Cloud, Developer & SRE teams. Leading cloud-native software engineering, platform reliability, observability, and operational excellence.
- Driving AI adoption across the software lifecycle — implementing AI-assisted development, testing, and automation across Cloud, Developer & SRE teams.
- Designed and built an AI-powered Problem Suggestion Engine using RAG architecture — deployed Ollama on AWS EKS serving Gemma 3 (LLM) and Nomic Embed Text (embedding model), ingesting runbooks into OpenSearch vector database, generating structured remediation suggestions (root cause analysis, troubleshooting commands, mitigation steps) as a POC for AIOps-driven incident resolution.
- Built AI-driven NiFi Flow Development tooling using Generative AI (Kiro IDE + Amazon Q) with MCP integrations to Confluence, Jira, and GitLab — enabling automatic generation of complete, deployable NiFi flows as Python scripts (NiPyApi), achieving 80% overall accuracy with human-in-the-loop validation, reducing flow development time from days to hours.
- Implemented Post Deployment Validation using AI agents that automatically verify deployment health, detect anomalies through pattern recognition, and flag regressions post-release.
- Designed Development Approval Process using AI to intelligently assess code changes against compliance policies, auto-generate review summaries, and streamline approval workflows.
- Led regular Kubernetes cluster upgrades (EKS) — managing control plane and node group version upgrades, add-on compatibility validation, PodDisruptionBudget enforcement, and rolling updates ensuring zero downtime and compliance with security patches.
- Performed proactive Kubernetes security patching including CVE remediation, container image vulnerability scanning, runtime security hardening, and CIS benchmark compliance across production clusters.
- Architected and deployed Apache NiFi clusters on AWS EKS using custom Helm charts, ensuring scalability, high availability, and automated failover/recovery.
- Built end-to-end high-volume communication infrastructure capable of handling 10M+ communications daily across multiple channels — Push notifications via IMI, high-volume Email through AWS SES (engineered to handle 3M+ emails/day), and SMS at enterprise scale.
- Leading platform reliability and observability using Dynatrace (full-stack APM, distributed tracing) and Prometheus/Grafana — ensuring platform health, performance, and service reliability.
- Driving operational excellence through proactive monitoring, vulnerability management, capacity planning, cloud platform utilisation cost reductions, and service performance optimisation.
- Leading resilience engineering, incident management, root cause analysis, and continuous improvement initiatives across the platform.
- Automated provisioning and deployments with GitLab CI/CD, Terraform, and ArgoCD — implementing versioned, governed release pipelines with rollback strategies.
- Integrated messaging systems including AWS SQS and Apache Kafka to support distributed, event-driven workflows and ensure reliable message delivery across microservices.
- Authored and deployed Python-based automation scripts hosted on AWS Lambda for diverse operational use cases, enhancing platform extensibility.
Senior DevOps Engineer
Informatica
Feb 2023 – Nov 2023
Bengaluru, Karnataka
Managed multi-cloud Kubernetes infrastructure and CI/CD platforms, driving deployment velocity and security posture improvements.
- Managed Kubernetes clusters on EKS & AKS, deploying microservices via Helm charts and custom Docker containers with production-grade configurations.
- Designed CI/CD pipelines using GitLab CI reducing deployment times by 80% with automated testing, artifact management, and release governance.
- Implemented HA deployments with Istio service mesh, blue/green deployment strategy, and canary rollouts for zero-downtime releases.
- Automated operational tasks with Python and Bash scripts, streamlining release processes and reducing manual intervention.
- Developed Infrastructure as Code with Terraform for cloud resource provisioning across AWS and Azure — managing state, modules, and drift detection.
- Configured full-stack observability with Prometheus, Grafana, and ELK Stack — custom dashboards, alerting rules, and distributed tracing for microservices.
- Integrated StackRox container security scans into CI/CD pipelines for proactive vulnerability detection, image signing, and runtime threat prevention (DevSecOps).
- Performed Kubernetes cluster upgrades, node scaling, and security patching across multi-cloud environments (EKS, AKS).
Cloud Engineer
Zoho Corporation
Feb 2018 – Feb 2023
Chennai, Tamil Nadu
Built and maintained large-scale on-premise and cloud infrastructure, managing 10,000+ servers with focus on automation, security, and reliability.
- Automated OS provisioning and application deployments with Cobbler, Puppet, and Ansible across 10,000+ servers — reducing provisioning time from days to minutes.
- Managed Kubernetes clusters (on-prem & AWS EKS) and implemented CI/CD pipelines with ECS and CodePipeline for containerised workloads.
- Administered core Linux services (DNS, DHCP, NTP, SMTP, Proxy) ensuring secure, reliable infrastructure across multiple datacenters.
- Configured HA load balancers (KeepAlived, Nginx) and database clusters (MySQL, PostgreSQL, Redis) with automated failover.
- Performed security patching, vulnerability remediation, and compliance hardening (CIS, STIG) across production environments.
- Managed data backup, disaster recovery procedures, and hardware failure resolution for on-premises infrastructure — maintaining RTO/RPO SLAs.
- Wrote Bash and Python scripts for server validation, enabling thorough pre-provisioning quality assurance testing across infrastructure fleet.
- Troubleshot diverse OS and hardware issues, identifying root causes and implementing effective solutions to maintain system stability and performance.
System Engineer
Pingserv Solutions LLC
May 2016 – Jan 2018
Chennai, Tamil Nadu
Managed Linux server fleet in remote datacenters with 99.99% uptime SLA, handling virtualisation, security, and infrastructure operations.
- Managed Linux servers in remote datacenters ensuring 99.99% uptime SLA compliance through proactive monitoring and rapid incident response.
- Configured virtualization platforms (XEN, KVM, OpenVZ) and administered core services (DNS, DHCP, LAMP stack) for hosting clients.
- Performed system monitoring, RAID configuration, and security audits — implementing hardening measures across server fleet.
- Executed server recovery operations including rescue mode, single-user mode, and emergency mode troubleshooting for critical production systems.
- Implemented anti-spam and anti-spoofing measures (SPF, DKIM, DMARC) to secure mail infrastructure.
- Managed file systems, LVM partitioning, disk quota administration, and storage capacity planning.
Certifications
CKA
Certified Kubernetes Administrator
CNCF
CKS
Certified Kubernetes Security Specialist
CNCF
Terraform Associate
HashiCorp Certified: Terraform Associate
HashiCorp
AWS AI Practitioner
AWS Certified AI Practitioner
Amazon Web Services
IBM RAG & Agentic AI
IBM RAG and Agentic AI Professional Certificate
IBM
RHCSA
Red Hat Certified System Administrator
Red Hat
RHCE
Red Hat Certified Engineer
Red Hat
ITIL Foundation
ITIL Foundation Certificate
Axelos
Get In Touch
I'm always open to discussing new opportunities, challenging projects, or how I can help build and scale your cloud infrastructure. Let's connect.
Location
Bengaluru, Karnataka, India