Open to Senior DevOps, SRE & AI Platform Roles

Hello, I'm

Numan Iftikhar

Senior |

I'm a Senior DevOps and AI Platform Engineer working across Azure, AWS, and Google Cloud. I design Kubernetes platforms, ship them with Terraform, and build the LLM systems on top — for teams that need production software to be reliable, secure, and defensible.

99.9% Uptime
70% Cost Saved
7 Certifications
5+ Years Exp
Numan Iftikhar - Senior DevOps, MLOps & AI Platform Engineer
stack: AWS Azure GCP K8s
Scroll Down

Trusted Technologies

AWS
Azure
GCP
Docker
Kubernetes
Terraform
Jenkins
GitHub Actions
FastAPI
Claude API

About Me

Senior DevOps, MLOps & AI Platform Engineer

I'm a Senior DevOps and AI Platform Engineer working across Azure, AWS, and Google Cloud. I design Kubernetes platforms, ship them with Terraform, and build the LLM systems on top — for teams that need production software to be reliable, secure, and defensible.

With 5+ years of experience across Azure (primary), AWS, and Google Cloud Platform, I am a Kubernetes Architect and IaC Expert with deep expertise in designing, deploying, and managing enterprise-grade cloud infrastructure across hybrid and multi-cloud environments — delivered in regulated environments spanning PCI-DSS, SOC 2, HIPAA, and ISO 27001.

I've extended that same production discipline into AI systems — RAG pipelines, local and hosted LLM serving, and agentic workflows built to run in production, not just demo — and into the data engineering plumbing (ingestion, embedding, API services) that makes models and reporting actually usable. I run this practice through my consultancy, 108 Core Technologies, and I'm currently open to senior DevOps, SRE, and AI Platform roles.

0+ Years Experience
0 Certifications
0 Cloud Platforms
0.9% Uptime Achieved
Name Numan Iftikhar
Email me@numaniftikhar.com
Phone +92 301 000 7414
Location Lahore, Pakistan (working worldwide)
Consultancy 108 Core Technologies
Languages English, Urdu
Availability Open to Senior Roles

Practice Areas

A multi-cloud infrastructure foundation, extended into AI systems and the data engineering that supports them

Cloud & DevOps Engineering

Production Kubernetes on AKS, infrastructure as code with Terraform, and CI/CD pipelines through Azure DevOps and GitHub Actions. Observability with Prometheus and Grafana. Delivered in regulated environments (PCI-DSS, SOC 2, HIPAA, ISO 27001).

AKSTerraformAzure DevOps PrometheusGrafana

AI Engineering & LLMOps

RAG pipelines, local and hosted LLM serving, and agentic workflows. Built to run in production, not just demo.

OllamaQdrantChromaDB FastAPILangGraphClaude API

Data Engineering

Ingestion, embedding, and API services in Python — the plumbing that makes models and reporting actually usable.

PythonFastAPIETL PipelinesCloud Pricing APIs

Currently building depth in: vLLM & model serving, fine-tuning (LoRA/PEFT), GPU scheduling, distributed training, and LLM eval/observability tooling (Langfuse, RAGAS/DeepEval) — hands-on learning, not yet production experience.

The Production Layer AI Needs

I bring the production layer AI needs — eval-gated CI/CD, LLM observability (token/cost/latency tracing), guardrails, autoscaling for inference, and FinOps — on top of a strong Kubernetes/Terraform foundation.

What I Can Do For You

End-to-end DevOps & Cloud solutions tailored to your business needs

Cloud Migration & Architecture

Migrate your workloads to AWS, Azure, or GCP with zero downtime. Design multi-cloud architectures optimized for cost, performance, and resilience.

  • Cloud readiness assessment
  • Multi-cloud strategy & design
  • Zero-downtime migration
  • Cost optimization (FinOps)

Kubernetes & Container Orchestration

Production-grade Kubernetes clusters on AKS, EKS, or GKE with auto-scaling, self-healing, and service mesh architectures.

  • Cluster setup & hardening
  • Helm charts & Istio mesh
  • Auto-scaling & spot nodes
  • 99.9%+ uptime guarantee

CI/CD Pipeline Engineering

Automated pipelines from commit to production using GitOps principles. Canary, blue/green, and rolling deployments.

  • GitHub Actions / GitLab CI / Jenkins
  • ArgoCD GitOps workflows
  • Automated testing gates
  • Container image scanning

Infrastructure as Code (IaC)

Reproducible, version-controlled infrastructure with Terraform and Ansible. Eliminate drift and manual provisioning forever.

  • Terraform modules & state mgmt
  • Ansible playbooks & roles
  • Multi-env provisioning
  • Drift detection & remediation

DevSecOps & Compliance

Security baked into every stage of your pipeline. Achieve HIPAA, PCI-DSS, SOC 2, and ISO 27001 compliance.

  • Vulnerability scanning (Trivy/Prisma)
  • Secrets management (Vault/KV)
  • Policy-as-code (OPA)
  • Compliance audit automation

Monitoring & Observability

Full-stack observability with Prometheus, Grafana, and cloud-native monitoring. Never be blind-sided by outages again.

  • Prometheus + Grafana dashboards
  • Log aggregation & alerting
  • SLO/SLI tracking
  • Incident response automation

How I Work

01

Discovery Call

Understand your infrastructure, pain points, and goals

02

Architecture Plan

Design a solution with diagrams, timelines, and milestones

03

Implementation

Build, test, and deploy with full transparency via Git

04

Handoff & Support

Documentation, knowledge transfer, and ongoing support

Professional Experience

March 2023 – Present

Senior DevOps Engineer

TrueMedIT Calgary, Alberta, Canada (Remote)
  • Multi-Cloud Architecture & Governance: Architected and managed enterprise infrastructure across Azure and GCP, implementing governance policies and cost-management strategies for high availability and cost-efficiency.
  • Advanced CI/CD Automation: Engineered end-to-end CI/CD pipelines using Jenkins, Azure DevOps, and Google Cloud Build, significantly reducing lead time for changes across multi-cloud environments.
  • Kubernetes Orchestration (AKS, GKE & Hybrid): Deployed and managed production-grade Kubernetes clusters ensuring 99.9% uptime through advanced scaling and self-healing configurations.
  • Enterprise IaC: Standardized multi-cloud provisioning using Terraform and Ansible, ensuring 100% consistency across Dev, QA, and Production stages.
  • DevSecOps & Compliance: Pioneered security integration with SSL/TLS, automated vulnerability scanning (Prisma/Trivy), and secrets management ensuring banking-level compliance.
  • Cloud Monitoring & Observability: Implemented centralized monitoring using Prometheus, Grafana, and Google Cloud Monitoring with integrated alerting for proactive incident detection.
  • Disaster Recovery: Designed automated backup and recovery solutions using Azure Backup, Velero, and GCP Cloud Storage with cross-region replication.
AzureGCPKubernetes TerraformJenkinsPrometheus
July 2021 – March 2023

DevOps Engineer

403 IT Solutions Texas, United States (Remote)
  • Hybrid Cloud Infrastructure: Architected and optimized Azure and GCP cloud environments, implementing governance and cost-management strategies that scaled infrastructure across multi-cloud deployments.
  • Enterprise Kubernetes & Virtualization: Designed high-availability Proxmox clusters on baremetal and orchestrated production Kubernetes workloads on AKS and GKE, ensuring 99.9% uptime.
  • Advanced CI/CD Orchestration: Engineered multi-platform automation pipelines using Jenkins, GitHub Actions, GitLab CI, and Google Cloud Build, reducing deployment lead times.
  • Strategic IaC: Automated end-to-end lifecycle of global environments across Azure and GCP using Terraform and Ansible, eliminating configuration drift.
  • Tier 3 Technical Leadership: Served as the final escalation point for complex networking and system architecture bottlenecks, resolving high-priority incidents.
  • Cross-functional Mentorship: Championed DevOps best practices, mentoring junior engineers and leading knowledge-sharing sessions.
AzureGCPProxmox AnsibleGitHub ActionsGitLab CI

Certifications

Professional Cloud DevOps Engineer

Google Cloud

Associate Cloud Engineer

Google Cloud

Professional Cloud Architect

Google Cloud

Solutions Architect – Associate

Amazon Web Services

Developer – Associate

Amazon Web Services

Azure Administrator Associate (AZ-104)

Microsoft

Terraform Associate (003)

HashiCorp

Featured Projects

Private RAG Pipeline Flagship

AI / LLMOps — Retrieval-Augmented Generation for Data That Can't Leave the Perimeter

OllamaQdrantFastAPIClaude API
  • Built a retrieval-augmented generation stack for teams whose data can't leave their perimeter, with local Ollama inference and a Qdrant vector store.
  • Served through a FastAPI service layer with clear ingestion, retrieval, and generation boundaries, and the Claude API wired in as an optional escalation path.
  • Dockerized for Kubernetes, applying the same production discipline as my infra work — containerized services, not a notebook prototype.
RAG LLM Dockerized

Cloud Cost Estimator

FinOps — LLM-Assisted Architecture Pricing

PythonAzure Pricing APILLM Tooling
  • An LLM interprets an architecture description and identifies the resources it implies, then Python prices them against the live Azure Retail Prices API.
  • Language models for reasoning, code for accuracy — the LLM never touches the arithmetic.
LLM FinOps

FinGuard Reference Architecture

Multi-Cloud Banking Platform

FinGuard Architecture Diagram Architecture Diagram
AWS EKSGCP GKETerraform ArgoCDIstioGitHub Actions VaultPrometheusOPA Gatekeeper
  • Architected production-grade multi-tenant banking infrastructure across AWS (primary) and GCP (DR), with automated cross-cloud failover achieving RPO < 5 min and RTO < 15 min.
  • Built full GitOps pipeline (GitHub Actions → ArgoCD) with canary deployments (10% → 50% → 100%), automated Trivy/Semgrep security scanning, and Cosign image signing — zero manual deployments to production.
  • Implemented zero-trust networking via Istio service mesh (mTLS), OPA Gatekeeper policy-as-code, and HashiCorp Vault for dynamic DB credential rotation every 24 hours.
  • Enforced PCI-DSS and SOC 2 compliance through automated policy gates, immutable audit trails, and per-tenant PostgreSQL database isolation via Terraform.
PCI-DSS SOC 2 Multi-Cloud GitOps

Lab Management System Reference Architecture

Multi-Tenant Healthcare SaaS on AKS

Lab Management System Architecture Diagram Architecture Diagram
Azure AKS.NET Core 8Helm IstioKEDAAzure DevOps Azure Key VaultTerraform
  • Designed HIPAA & ISO 27001-compliant multi-tenant SaaS platform for hospitals and research labs, with full isolation at compute, data, and identity layers.
  • Implemented KEDA-driven autoscaling (0 → 20 pods) triggered by Azure Service Bus queue depth, running on Spot nodes — achieving ~70% cost reduction.
  • Delivered p99 API response time < 100ms and 99.97% uptime SLA via AKS with HPA, PodDisruptionBudgets, and Azure Front Door.
  • Provisioned full infrastructure with Terraform (AKS, Azure SQL, Redis, Cosmos DB, Blob Storage) and deployed via Helm + ArgoCD GitOps pipeline with blue/green canary strategy.
  • The same event-driven, scale-to-zero pattern (KEDA + Spot nodes) that inference workloads need — elastic capacity that follows real demand instead of idling on fixed nodes.
HIPAA ISO 27001 <100ms p99 99.97%

Secure Hybrid VoIP Platform

Telecom — Hybrid Cloud Infrastructure & VoIP (Dubai)

Azure AKSTerraformAzure DevOps Site-to-Site IPSec VPNFreeSwitch VoIP NSGsK8s NetworkPolicies
  • Architected production AKS infrastructure connected to on-premises systems in Dubai via an Azure Site-to-Site IPSec VPN, provisioned entirely with Terraform and Azure DevOps.
  • Deployed and operated a FreeSwitch VoIP stack alongside the Kubernetes workloads over the hybrid link.
  • Enforced defense-in-depth network segmentation with NSGs at the Azure network layer and Kubernetes NetworkPolicies at the pod layer.
Hybrid Cloud Network Isolation IaC

108 Core Technologies

Consultancy

  • My consultancy — cloud and AI engineering for SMEs, with direct client ownership and no handoffs.
Cloud AI Engineering

In Progress / Roadmap

What I'm actively building depth in next — not yet shipped to production

Learning

LLM Serving & Fine-Tuning Lab

Hands-on lab exploring vLLM for model serving and fine-tuning with LoRA/PEFT.

Learning

GPU Scheduling & Distributed Training

Deepening my Kubernetes background into GPU-aware scheduling and distributed training workloads.

Why Work With Me

100% Remote-Ready

5+ years working remotely with teams in Canada, USA, and worldwide. Async-first, timezone-flexible, and self-managed.

7 Cloud Certifications

Certified across all 3 major clouds (AWS, Azure, GCP) plus Terraform. I don't just talk the talk.

Proven Results

99.9% uptime, 70% cost reductions, zero-downtime deployments. My work is measured in business outcomes.

Security-First Mindset

PCI-DSS, HIPAA, SOC 2, ISO 27001 — I build compliance into the infrastructure from day one.

Clear Communication

Regular updates, documented decisions, and no jargon. You'll always know exactly where your project stands.

Long-Term Partner

I don't just deploy and disappear. I provide knowledge transfer, documentation, and ongoing support.

Ready to Level Up Your Infrastructure?

Whether you need a full-time remote DevOps engineer or a freelance cloud architect for your next project — let's build something great together.

What Clients Say

"Numan transformed our entire cloud infrastructure. Migrated us from a single-server setup to a fully automated multi-cloud architecture with zero downtime. Our deployment time went from hours to minutes."

Healthcare SaaS Client CTO, TrueMedIT

"Exceptional Kubernetes expertise. Numan set up our production clusters with auto-scaling and monitoring that just works. We haven't had a single unplanned outage since he built our infrastructure."

Enterprise IT Client VP Engineering, 403 IT Solutions

"Numan's DevSecOps implementation saved us from a potential compliance nightmare. He automated our entire security pipeline and got us PCI-DSS certified ahead of schedule. Highly recommended."

FinTech Startup Client Founder & CEO

Want to be my next success story?

Book a Free Call

Education

Bachelor of Science in Information Technology

Virtual University of Pakistan May 2018 – September 2022 CGPA: 3.2 / 4.0

Latest Insights

Sharing knowledge on DevOps, Cloud Architecture, and Infrastructure best practices

Multi-Cloud
March 2026 8 min read

Multi-Cloud Strategy: AWS + GCP Failover Architecture That Achieves RPO < 5 Minutes

How I designed a production-grade cross-cloud disaster recovery system for a banking platform using Terraform, Istio, and automated failover pipelines.

AWSGCPTerraformDR
Read Article
Kubernetes
February 2026 10 min read

KEDA Autoscaling on AKS: How We Cut Cloud Costs by 70% for Bursty Workloads

A deep dive into event-driven autoscaling with KEDA, Azure Service Bus, and Spot nodes — from zero pods to handling 10K concurrent requests.

KubernetesKEDAAzureFinOps
Read Article
DevSecOps
January 2026 12 min read

Zero-Trust DevSecOps: Building PCI-DSS Compliant CI/CD Pipelines from Scratch

Step-by-step guide to implementing automated security scanning, OPA policy gates, Vault secret rotation, and Cosign image signing in your GitOps workflow.

DevSecOpsOPAVaultCI/CD
Read Article

Book a Free Consultation

Need help with cloud architecture, DevOps strategy, or infrastructure optimization? Let's talk!

30 Minutes

Free one-on-one consultation call

Google Meet / Zoom

Virtual meeting at your convenience

100% Free

No obligations, no hidden charges

What We Can Discuss

  • Cloud Migration Strategy (AWS, Azure, GCP)
  • Kubernetes Architecture & Deployment
  • CI/CD Pipeline Design & Optimization
  • Infrastructure as Code (Terraform/Ansible)
  • DevSecOps & Security Best Practices
  • Cloud Cost Optimization (FinOps)
  • Monitoring & Observability Setup

Powered by Calendly. Pick a time that works for you — confirmation is instant.

Get In Touch

Have a project in mind or want to discuss cloud infrastructure? Let's connect!

Location

Lahore, Pakistan (working worldwide)