HI, I'M

Prakhar Mathur

SITE RELIABILITY ENGINEER & AIOPS SPECIALIST

Building scalable, AI-ready infrastructure on AWS & Kubernetes. Specializing in AIOps, MLOps observability, automation, and cloud-native systems — keeping production fast, resilient, and self-healing.

CLOUD NATIVE

AWS, GCP, Azure

KUBERNETES

EKS, AKS, GKE

OBSERVABILITY

Prometheus, Grafana

AUTOMATION

Terraform, Ansible, CI/CD

ABOUT ME

Building reliable infrastructure that scales in production.

I'm a Site Reliability Engineer (SRE) and DevOps Engineer specializing in building scalable, resilient, and observable cloud infrastructure. My expertise lies in Kubernetes, AWS, CI/CD pipelines, infrastructure automation, and cloud-native systems designed for reliability at scale.

I work across production environments to improve system performance, availability, and operational efficiency through monitoring, automation, incident response, and infrastructure optimization.

My core stack includes Kubernetes, AWS, Terraform, Docker, ArgoCD, Prometheus, Grafana, ELK, and GitHub Actions. I'm also exploring AIOps, intelligent monitoring, and automation to reduce operational overhead.

3+

YEARS EXPERIENCE

SRE / DevOps Engineer

Cloud & On-Prem

INFRASTRUCTURE

Managing hybrid infrastructure across environments

99.9%

RELIABILITY FOCUS

Designing systems that stay fast and available

AI + SRE

EXPLORING AIOPS

Intelligent monitoring, automation & operations

WHAT I DO

Cloud Infrastructure

Building scalable cloud infrastructure using AWS, Kubernetes, Terraform, and cloud-native architecture.

AWSKubernetesTerraformDocker

Observability & Monitoring

Implementing monitoring, logging, and observability systems for production workloads.

PrometheusGrafanaELKNew Relic

Automation & DevOps

Automating deployments, CI/CD pipelines, and infrastructure workflows for faster delivery.

GitHub ActionsArgoCDBashCI/CD

AI + SRE Systems

Exploring AI-powered operations, intelligent alerting, and automation for modern SRE workflows.

AIOpsAlertingMLAutomation

CURRENTLY AT

IQM Corporation

Site Reliability Engineer

BASED IN

India

Open to remote opportunities

FOCUS AREAS

Kubernetes • Cloud • Observability

Automation • Reliability • AIOps

I BELIEVE IN

Systems that heal,

automate and scale.

SKILLS & EXPERTISE

SRE & DevOps Tech Stack — Kubernetes, AWS, Terraform & More.

A production-grade stack built around Site Reliability Engineering, cloud-native infrastructure, observability, automation, and AIOps.

Infrastructure & Cloud

Cloud platforms, container orchestration, and infrastructure automation at scale.

  • Linux
  • AWS
  • Azure
  • Docker
  • Kubernetes
  • Terraform
  • ArgoCD

Observability

Monitoring, logging, alerting, and full-stack observability for production systems.

  • Prometheus
  • Grafana
  • New Relic
  • Elasticsearch
  • Kibana
  • Robusta
  • DataDog

Languages & Dev

Programming languages and frontend technologies for tooling and automation.

  • Python
  • C++
  • JavaScript
  • ReactJS
  • TailwindCSS
  • HTML
  • CSS

Data & ML

Data engineering, machine learning frameworks, and database management.

  • Tensorflow
  • Keras
  • Scikit-Learn
  • Power BI
  • PostgreSQL
  • MySQL
  • Kafka

AIOps & Automation

AI-powered operations, intelligent alerting, and workflow automation.

  • AIOps
  • MLOps
  • Ansible
  • GitHub Actions
  • CI/CD
  • Bash

Tools & Platforms

Developer tools, project management, and collaboration platforms.

  • GitHub
  • Jira
  • Postman
  • Lens
  • Graylog
  • Argo CD

30+

Technologies

Across cloud, infra & dev

6

Core Categories

Infra, Obs, Dev, Data, AI, Tools

3+

Years Hands-on

Production-grade experience

Always

Learning

AIOps, MLOps & cloud-native

WORK EXPERIENCE

Site Reliability Engineer — Career Timeline.

3+ years building and operating production-grade cloud infrastructure, Kubernetes platforms, and observability systems.

Building AI-ready observability infrastructure for high-frequency algorithmic bidding engines.

  • Engineered AI-ready observability pipelines via DataDog & Grafana for real-time latency tracking
  • Implemented anomaly detection & behavioral analysis for performance deviation monitoring
  • Defined and monitored SLOs, SLIs, and SLAs in collaboration with Data and AdOps teams
  • Led capacity planning and performance optimization for data-intensive workloads

3+

Years Experience

SRE / DevOps Engineering

3

Companies

Production environments

60%

Faster Alerting

Reduced incident response time

24/7

Reliability Focus

Always-on production systems

PROJECTS

ML, Fullstack & Frontend Projects.

A selection of machine learning, deep learning, and fullstack projects spanning data science, computer vision, and modern web development.

House Price Prediction
ML / Data

House Price Prediction

Machine Learning · Data Analytics

ML regression model to predict house prices using feature engineering, Scikit-Learn, and data visualization pipelines.

  • Python
  • Scikit-Learn
  • Pandas
  • EDA
Customer Churn Prediction
ML / Data

Customer Churn Prediction

Machine Learning · Data Analytics

Telecom churn prediction for Reliance Jio using classification models with feature selection and business insight reporting.

  • Python
  • XGBoost
  • Feature Engineering
  • Visualization
Sign Language Detection
Deep Learning

Sign Language Detection

Machine Learning · Deep Learning

Real-time sign language recognition system using CNN and computer vision for accessibility and communication aid.

  • TensorFlow
  • Keras
  • OpenCV
  • CNN
Ecommerce Project
Frontend

Ecommerce Project

Frontend · Ecommerce

Fully responsive ecommerce storefront with product listing, cart, and modern UI built with React and TailwindCSS.

  • ReactJS
  • TailwindCSS
  • JavaScript
  • Netlify
Job Portal
Fullstack

Job Portal

Fullstack · Firebase

Full-stack job portal with Firebase authentication, real-time database, and role-based access for recruiters and applicants.

  • ReactJS
  • Firebase
  • Vercel
  • Auth
Fitclub Gym Website
Frontend

Fitclub Gym Website

Frontend

Modern gym landing page with animated hero section, service cards, and responsive layout built for conversion.

  • ReactJS
  • CSS
  • Animations
  • Netlify

6

Projects

ML, Fullstack & Frontend

3

ML Projects

Prediction & deep learning

React

Frontend Stack

ReactJS + TailwindCSS

GitHub

Open Source

All projects on GitHub

LATEST WRITING

From the Blog.

View all articles
site-reliability-engineeraws-sqs

AWS SQS Explained: The Complete Beginner’s Guide to Amazon Simple Queue Service

9 min · May 31, 2026Read
machine-learning-aicontext-engineering

Context Engineering is the real AI skill no one talks about

3 min · May 6, 2026Read
sreai-governance

Managing and Observing Kubernetes with MCP, AI-Governed, and Compliance-First

3 min · Feb 26, 2026Read

GET IN TOUCH

Let's Build Something Together.

Open to SRE, DevOps, and AIOps opportunities. Whether it's a full-time role, freelance project, or just a tech conversation — reach out.

Email

mathurprakhar1@gmail.com

Location

India — Available Worldwide

Remote-first · Open to global opportunities

Availability

Open to opportunities

SRE · DevOps · AIOps roles

Send a Message