Skip to content

AI/ML systems

The production-ML side of the AI stack: training infrastructure, model deployment, MLOps, feature stores, and the cloud platforms that orchestrate it. Where LLMs and GenAI covers language-model specifics, this page covers the broader systems engineering: classical ML, deep learning, GPU infrastructure, ML pipelines, and the certs that test them.

flowchart LR
  D[Data + features] --> TR[Training infra<br/>GPUs, distributed training]
  TR --> REG[Model registry]
  REG --> SRV[Serving:<br/>online / batch / edge]
  SRV --> MON[Monitoring + drift]
  MON -. retrain .-> TR

Learn


Compare


Reference


Build


Certify

Certs that test ML systems engineering specifically:

Foundational - AWS AI Practitioner - cross-cert GenAI study track - Azure AI Fundamentals (AI-900)

Associate - AWS ML Engineer (MLA-C01) - the production-ML cert - Azure AI Engineer (AI-102) - GCP Machine Learning Engineer - Databricks ML Associate - Databricks-flavored MLOps - Databricks GenAI Engineer Associate - NVIDIA AI Infrastructure & Operations Associate - GPU infra - NVIDIA GenAI/LLM Associate

Specialty / Professional - AWS Machine Learning Specialty (MLS-C01) - Databricks ML Professional - NVIDIA AI Infrastructure Professional - NVIDIA AI Operations Professional - NVIDIA Accelerated Data Science Professional


Roadmap

The career-track view: AI/ML Engineer roadmap.