Open to senior DevOps, SRE, and MLOps roles

Numan Iftikhar

Solution Architect for Cloud & AI Systems

Senior DevOps engineer, 5+ years. I design and run infrastructure on Azure, AWS, and GCP — Kubernetes, GitOps with Argo CD, and the data/LLM layer on top of that. Most of the work has been in regulated setups (PCI-DSS, HIPAA), not slide-deck architecture.

  • DevOps & SRE
  • Data engineering
  • MLOps
  • Azure · AWS · GCP

Lahore · Working with Unduit · me@numaniftikhar.com

Numan Iftikhar

Core Technologies & Data Stack

AWS
Azure
GCP
Docker
Kubernetes
Terraform
Apache Kafka
Apache Spark
Airflow
dbt Core
Snowflake
Delta Lake
PostgreSQL
GitHub Actions
FastAPI
Claude API

About Me

I Design, Automate & Scale Cloud Systems

I'm a Senior DevOps and AI Platform Engineer working across Azure, AWS, and Google Cloud. I design Kubernetes platforms, ship them with Terraform, and build the LLM systems on top — for teams that need production software to be reliable, secure, and defensible.

With 5+ years of experience across Azure (primary), AWS, and Google Cloud Platform, I am a Kubernetes Architect and IaC Expert with deep expertise in designing, deploying, and managing enterprise-grade cloud infrastructure across hybrid and multi-cloud environments — delivered in regulated environments spanning PCI-DSS, SOC 2, HIPAA, and ISO 27001.

I've extended that same production discipline into AI systems — RAG pipelines, local and hosted LLM serving, and agentic workflows built to run in production, not just demo — and into the data engineering plumbing (ingestion, embedding, API services) that makes models and reporting actually usable. I run this practice through my consultancy, 108 Core Technologies, and I'm currently open to senior DevOps, SRE, and AI Platform roles.

5+ Years Experience
7 Certifications
3 Cloud Platforms
99.9% Uptime Achieved
Name Numan Iftikhar
Email me@numaniftikhar.com
Phone +92 301 000 7414
Location Lahore, Pakistan (working worldwide)
Consultancy 108 Core Technologies
Languages English, Urdu
Availability Open to Senior Roles

Practice Areas

A multi-cloud infrastructure foundation, extended into real-time Data Engineering and production AI systems

Cloud & DevOps Engineering

Production Kubernetes on AKS/EKS/GKE, infrastructure as code with Terraform, and CI/CD pipelines through Azure DevOps and GitHub Actions. Observability with Prometheus and Grafana. Delivered in regulated environments (PCI-DSS, SOC 2, HIPAA, ISO 27001).

AKS / EKS / GKETerraformAzure DevOps GitHub ActionsPrometheusGrafana

Data Engineering & Modern Lakehouse

High-throughput real-time streaming and batch ETL/ELT pipelines, modern Lakehouse architectures (Delta Lake, Snowflake), automated data transformation with dbt, Airflow workflow orchestration, and high-dimensional Vector ETL for LLMs.

Apache KafkaPySparkAirflow dbt CoreDelta LakeSnowflakepgvector

AI Engineering & LLMOps

Enterprise RAG pipelines, local (Ollama/vLLM) and hosted LLM serving, and agentic workflows. Eval-gated deployments, guardrails, and token cost tracking built to run in production, not just demo.

OllamaQdrantChromaDB FastAPILangGraphClaude API

Currently building depth in: vLLM & model serving, fine-tuning (LoRA/PEFT), GPU scheduling, distributed training, and LLM eval/observability tooling (Langfuse, RAGAS/DeepEval) — hands-on learning, not yet production experience.

The Production Foundation Data & AI Need

I bridge the gap between infrastructure, high-throughput streaming data pipelines, and AI systems — combining Kafka/Spark Lakehouse plumbing with eval-gated LLMOps, Kubernetes scaling, and FinOps cost controls.

What I Can Do For You

End-to-end DevOps, Data Engineering & Cloud AI solutions tailored to your business needs

Cloud Migration & Architecture

Migrate your workloads to AWS, Azure, or GCP with zero downtime. Design multi-cloud architectures optimized for cost, performance, and resilience.

  • Cloud readiness assessment
  • Multi-cloud strategy & design
  • Zero-downtime migration
  • Cost optimization (FinOps)

Kubernetes & Container Orchestration

Production-grade Kubernetes clusters on AKS, EKS, or GKE with auto-scaling, self-healing, and service mesh architectures.

  • Cluster setup & hardening
  • Helm charts & Istio mesh
  • Auto-scaling & spot nodes
  • 99.9%+ uptime guarantee

Data Engineering & Lakehouses

Resilient streaming data pipelines and modern lakehouses. Automated ETL/ELT workflows, real-time CDC, and vector database indexing.

  • Apache Kafka & Spark streaming
  • Delta Lake & Snowflake modeling
  • dbt transformations & Airflow DAGs
  • Vector DB ingestion (Qdrant/pgvector)

CI/CD Pipeline Engineering

Automated pipelines from commit to production using GitOps principles. Canary, blue/green, and rolling deployments.

  • GitHub Actions / GitLab CI / Jenkins
  • ArgoCD GitOps workflows
  • Automated testing gates
  • Container image scanning

Infrastructure as Code (IaC)

Reproducible, version-controlled infrastructure with Terraform and Ansible. Eliminate drift and manual provisioning forever.

  • Terraform modules & state mgmt
  • Ansible playbooks & roles
  • Multi-env provisioning
  • Drift detection & remediation

DevSecOps & Compliance

Security baked into every stage of your pipeline. Achieve HIPAA, PCI-DSS, SOC 2, and ISO 27001 compliance.

  • Vulnerability scanning (Trivy/Prisma)
  • Secrets management (Vault/KV)
  • Policy-as-code (OPA)
  • Compliance audit automation

Monitoring & Observability

Full-stack observability with Prometheus, Grafana, and cloud-native monitoring. Never be blind-sided by outages again.

  • Prometheus + Grafana dashboards
  • Log aggregation & alerting
  • SLO/SLI tracking
  • Incident response automation

How I Work

01

Discovery Call

Understand your infrastructure, pain points, and goals

02

Architecture Plan

Design a solution with diagrams, timelines, and milestones

03

Implementation

Build, test, and deploy with full transparency via Git

04

Handoff & Support

Documentation, knowledge transfer, and ongoing support

Professional Experience

March 2023 – Present

Senior DevOps Engineer

TrueMedIT Calgary, Alberta, Canada (Remote)
  • Multi-Cloud Architecture & Governance: Architected and managed enterprise infrastructure across Azure and GCP, implementing governance policies and cost-management strategies for high availability and cost-efficiency.
  • Advanced CI/CD Automation: Engineered end-to-end CI/CD pipelines using Jenkins, Azure DevOps, and Google Cloud Build, significantly reducing lead time for changes across multi-cloud environments.
  • Kubernetes Orchestration (AKS, GKE & Hybrid): Deployed and managed production-grade Kubernetes clusters ensuring 99.9% uptime through advanced scaling and self-healing configurations.
  • Enterprise IaC: Standardized multi-cloud provisioning using Terraform and Ansible, ensuring 100% consistency across Dev, QA, and Production stages.
  • DevSecOps & Compliance: Pioneered security integration with SSL/TLS, automated vulnerability scanning (Prisma/Trivy), and secrets management ensuring banking-level compliance.
  • Cloud Monitoring & Observability: Implemented centralized monitoring using Prometheus, Grafana, and Google Cloud Monitoring with integrated alerting for proactive incident detection.
  • Data Engineering & Streaming Pipelines: Engineered real-time event streaming and batch ETL pipelines using Apache Kafka, PySpark, and dbt Core across Snowflake and PostgreSQL, ensuring automated schema evolution and data quality enforcement.
  • Disaster Recovery: Designed automated backup and recovery solutions using Azure Backup, Velero, and GCP Cloud Storage with cross-region replication.
AzureGCPKubernetes KafkaPySparkTerraformJenkins
July 2021 – March 2023

DevOps & Cloud Engineer

403 IT Solutions Texas, United States (Remote)
  • Hybrid Cloud Infrastructure: Architected and optimized Azure and GCP cloud environments, implementing governance and cost-management strategies that scaled infrastructure across multi-cloud deployments.
  • Enterprise Kubernetes & Virtualization: Designed high-availability Proxmox clusters on baremetal and orchestrated production Kubernetes workloads on AKS and GKE, ensuring 99.9% uptime.
  • Data Ingestion & Automation: Automated database migrations, telemetry stream collectors, and multi-stage ETL scripts in Python for analytical dashboards.
  • Advanced CI/CD Orchestration: Engineered multi-platform automation pipelines using Jenkins, GitHub Actions, GitLab CI, and Google Cloud Build, reducing deployment lead times.
  • Strategic IaC: Automated end-to-end lifecycle of global environments across Azure and GCP using Terraform and Ansible, eliminating configuration drift.
  • Tier 3 Technical Leadership: Served as the final escalation point for complex networking and system architecture bottlenecks, resolving high-priority incidents.
  • Cross-functional Mentorship: Championed DevOps best practices, mentoring junior engineers and leading knowledge-sharing sessions.
AzureGCPKubernetesPython AnsibleGitHub ActionsGitLab CI

Certifications

Professional Cloud DevOps Engineer

Google Cloud

Associate Cloud Engineer

Google Cloud

Professional Cloud Architect

Google Cloud

Solutions Architect – Associate

Amazon Web Services

Developer – Associate

Amazon Web Services

Azure Administrator Associate (AZ-104)

Microsoft

Terraform Associate (003)

HashiCorp

Featured Projects

Private RAG Pipeline Flagship

AI / LLMOps — Retrieval-Augmented Generation for Data That Can't Leave the Perimeter

OllamaQdrantFastAPIClaude API
  • Built a retrieval-augmented generation stack for teams whose data can't leave their perimeter, with local Ollama inference and a Qdrant vector store.
  • Served through a FastAPI service layer with clear ingestion, retrieval, and generation boundaries, and the Claude API wired in as an optional escalation path.
  • Dockerized for Kubernetes, applying the same production discipline as my infra work — containerized services, not a notebook prototype.
RAG LLM Dockerized

Real-Time Streaming & Lakehouse Platform Data Eng

Data Engineering — High-Throughput Kafka + Spark Stream Processing & Delta Lake Medallion Architecture

Apache KafkaPySparkDelta Lakedbt CoreAirflowpgvector
  • Designed and deployed an event-driven data streaming engine processing 50M+ events/day via Apache Kafka clusters and distributed PySpark Streaming jobs on Kubernetes.
  • Built a Medallion Lakehouse (Bronze → Silver → Gold) on Delta Lake with automated dbt Core SQL transformations, schema enforcement, and Great Expectations quality assertions.
  • Integrated automated Vector & Metadata ETL pipelines that continuously stream clean text chunks into Qdrant and pgvector for sub-second semantic retrieval across enterprise data stores.
  • Implemented end-to-end Airflow DAG orchestration with SLA monitoring, automated backfills, and Prometheus lag metrics — achieving 99.98% pipeline reliability.
Delta Lake Kafka Streaming dbt & Airflow Vector ETL

Cloud Cost Estimator

FinOps — LLM-Assisted Architecture Pricing

PythonAzure Pricing APILLM Tooling
  • An LLM interprets an architecture description and identifies the resources it implies, then Python prices them against the live Azure Retail Prices API.
  • Language models for reasoning, code for accuracy — the LLM never touches the arithmetic.
LLM FinOps

FinGuard Reference Architecture

Multi-Cloud Banking Platform

FinGuard Architecture Diagram Architecture Diagram
AWS EKSGCP GKETerraform ArgoCDIstioGitHub Actions VaultPrometheusOPA Gatekeeper
  • Architected production-grade multi-tenant banking infrastructure across AWS (primary) and GCP (DR), with automated cross-cloud failover achieving RPO < 5 min and RTO < 15 min.
  • Built full GitOps pipeline (GitHub Actions → ArgoCD) with canary deployments (10% → 50% → 100%), automated Trivy/Semgrep security scanning, and Cosign image signing — zero manual deployments to production.
  • Implemented zero-trust networking via Istio service mesh (mTLS), OPA Gatekeeper policy-as-code, and HashiCorp Vault for dynamic DB credential rotation every 24 hours.
  • Enforced PCI-DSS and SOC 2 compliance through automated policy gates, immutable audit trails, and per-tenant PostgreSQL database isolation via Terraform.
PCI-DSS SOC 2 Multi-Cloud GitOps

Lab Management System Reference Architecture

Multi-Tenant Healthcare SaaS on AKS

Lab Management System Architecture Diagram Architecture Diagram
Azure AKS.NET Core 8Helm IstioKEDAAzure DevOps Azure Key VaultTerraform
  • Designed HIPAA & ISO 27001-compliant multi-tenant SaaS platform for hospitals and research labs, with full isolation at compute, data, and identity layers.
  • Implemented KEDA-driven autoscaling (0 → 20 pods) triggered by Azure Service Bus queue depth, running on Spot nodes — achieving ~70% cost reduction.
  • Delivered p99 API response time < 100ms and 99.97% uptime SLA via AKS with HPA, PodDisruptionBudgets, and Azure Front Door.
  • Provisioned full infrastructure with Terraform (AKS, Azure SQL, Redis, Cosmos DB, Blob Storage) and deployed via Helm + ArgoCD GitOps pipeline with blue/green canary strategy.
  • The same event-driven, scale-to-zero pattern (KEDA + Spot nodes) that inference workloads need — elastic capacity that follows real demand instead of idling on fixed nodes.
HIPAA ISO 27001 <100ms p99 99.97%

Secure Hybrid VoIP Platform

Telecom — Hybrid Cloud Infrastructure & VoIP (Dubai)

Azure AKSTerraformAzure DevOps Site-to-Site IPSec VPNFreeSwitch VoIP NSGsK8s NetworkPolicies
  • Architected production AKS infrastructure connected to on-premises systems in Dubai via an Azure Site-to-Site IPSec VPN, provisioned entirely with Terraform and Azure DevOps.
  • Deployed and operated a FreeSwitch VoIP stack alongside the Kubernetes workloads over the hybrid link.
  • Enforced defense-in-depth network segmentation with NSGs at the Azure network layer and Kubernetes NetworkPolicies at the pod layer.
Hybrid Cloud Network Isolation IaC

108 Core Technologies

Consultancy

  • My consultancy — cloud and AI engineering for SMEs, with direct client ownership and no handoffs.
Cloud AI Engineering

In Progress / Roadmap

What I'm actively building depth in next — not yet shipped to production

Learning

LLM Serving & Fine-Tuning Lab

Hands-on lab exploring vLLM for model serving and fine-tuning with LoRA/PEFT.

Learning

GPU Scheduling & Distributed Training

Deepening my Kubernetes background into GPU-aware scheduling and distributed training workloads.

Why Work With Me

100% Remote-Ready

5+ years working remotely with teams in Canada, USA, and worldwide. Async-first, timezone-flexible, and self-managed.

7 Cloud Certifications

Certified across all 3 major clouds (AWS, Azure, GCP) plus Terraform. I don't just talk the talk.

Proven Results

99.9% uptime, 70% cost reductions, zero-downtime deployments. My work is measured in business outcomes.

Security-First Mindset

PCI-DSS, HIPAA, SOC 2, ISO 27001 — I build compliance into the infrastructure from day one.

Clear Communication

Regular updates, documented decisions, and no jargon. You'll always know exactly where your project stands.

Long-Term Partner

I don't just deploy and disappear. I provide knowledge transfer, documentation, and ongoing support.

Ready to Level Up Your Infrastructure?

Whether you need a full-time remote DevOps engineer or a freelance cloud architect for your next project — let's build something great together.

What Clients Say

"Numan transformed our entire cloud infrastructure. Migrated us from a single-server setup to a fully automated multi-cloud architecture with zero downtime. Our deployment time went from hours to minutes."

Healthcare SaaS Client CTO, TrueMedIT

"Exceptional Kubernetes expertise. Numan set up our production clusters with auto-scaling and monitoring that just works. We haven't had a single unplanned outage since he built our infrastructure."

Enterprise IT Client VP Engineering, 403 IT Solutions

"Numan's DevSecOps implementation saved us from a potential compliance nightmare. He automated our entire security pipeline and got us PCI-DSS certified ahead of schedule. Highly recommended."

FinTech Startup Client Founder & CEO

Want to be my next success story?

Book a Free Call

Education

Bachelor of Science in Information Technology

Virtual University of Pakistan May 2018 – September 2022 CGPA: 3.2 / 4.0

Latest Insights

Sharing knowledge on DevOps, Cloud Architecture, and Infrastructure best practices

Multi-Cloud
March 2026 8 min read

Multi-Cloud Strategy: AWS + GCP Failover Architecture That Achieves RPO < 5 Minutes

How I designed a production-grade cross-cloud disaster recovery system for a banking platform using Terraform, Istio, and automated failover pipelines.

AWSGCPTerraformDR
Read Article
Kubernetes
February 2026 10 min read

KEDA Autoscaling on AKS: How We Cut Cloud Costs by 70% for Bursty Workloads

A deep dive into event-driven autoscaling with KEDA, Azure Service Bus, and Spot nodes — from zero pods to handling 10K concurrent requests.

KubernetesKEDAAzureFinOps
Read Article
DevSecOps
January 2026 12 min read

Zero-Trust DevSecOps: Building PCI-DSS Compliant CI/CD Pipelines from Scratch

Step-by-step guide to implementing automated security scanning, OPA policy gates, Vault secret rotation, and Cosign image signing in your GitOps workflow.

DevSecOpsOPAVaultCI/CD
Read Article

Book a Free Consultation

Need help with cloud architecture, DevOps strategy, or infrastructure optimization? Let's talk!

30 Minutes

Free one-on-one consultation call

Google Meet / Zoom

Virtual meeting at your convenience

100% Free

No obligations, no hidden charges

What We Can Discuss

  • Cloud Migration Strategy (AWS, Azure, GCP)
  • Kubernetes Architecture & Deployment
  • Real-Time Data Streaming (Kafka & Spark)
  • Lakehouse Architecture & dbt Modeling
  • Vector DB & Metadata ETL for Enterprise RAG
  • CI/CD Pipeline Design & Optimization
  • Infrastructure as Code (Terraform/Ansible)
  • DevSecOps & Security Compliance (PCI, SOC 2, HIPAA)
  • Cloud Cost Optimization (FinOps)

Powered by Calendly. Pick a time that works for you — confirmation is instant.

Get In Touch

Have a project in mind or want to discuss cloud infrastructure? Let's connect!

Location

Lahore, Pakistan (working worldwide)