Skip to content

Microsoft Fabric Analytics Engineer (DP-600) - Fact Sheet

Quick Reference

Exam Code: DP-600 Duration: 120 minutes Questions: 40-60 questions Passing Score: 700/1000 Cost: $165 USD Validity: 3 years Difficulty: ⭐⭐⭐⭐

Exam Domains

Domain Weight Key Focus
Plan, implement, and manage a solution for data analytics 10-15% Fabric workspace, capacity, governance
Prepare and serve data 40-45% Lakehouses, data pipelines, semantic models
Implement and manage semantic models 20-25% Data modeling, DAX, optimization
Explore and analyze data 20-25% Power BI reports, KQL queries, notebooks

Microsoft Fabric Overview

What is Microsoft Fabric? - Unified analytics platform (SaaS) - Combines Data Engineering, Data Science, Data Warehousing, Real-time Analytics, Power BI - Built on OneLake (unified data lake) - Single capacity-based pricing model - πŸ“– Microsoft Fabric Documentation - Complete Fabric guide - πŸ“– Fabric Get Started - Platform overview - πŸ“– Fabric Architecture - Technical architecture - πŸ“– Fabric Licensing - Capacity and licensing

OneLake

OneLake: Unified Data Lake - Single, hierarchical namespace - ADLS Gen2 compatible - Automatic data organization - Delta Lake format by default - Shortcuts: Access data without copying - πŸ“– OneLake Overview - OneLake concepts - πŸ“– OneLake Data Hub - Data discovery - πŸ“– OneLake Shortcuts - Data federation - πŸ“– OneLake Security - Access control - πŸ“– OneLake Integration - Azure Storage integration

Fabric Workspaces

Workspaces - Containers for Fabric items - Role-based access: Admin, Member, Contributor, Viewer - Workspace capacity assignment - Licensing modes: Trial, Premium, Fabric - πŸ“– Workspaces - Workspace management - πŸ“– Workspace Roles - Permission levels - πŸ“– Workspace Identity - Service principal access

Lakehouses

Lakehouse Architecture - Combines data lake (Parquet) and warehouse (SQL endpoint) - Delta Lake format for ACID transactions - Automatic metadata generation - SQL analytics endpoint (read-only) - Notebooks for data engineering - πŸ“– Lakehouse Overview - Lakehouse concepts - πŸ“– Create Lakehouse - Setup guide - πŸ“– Lakehouse SQL Endpoint - SQL analytics - πŸ“– Tables in Lakehouse - Table management - πŸ“– Lakehouse Files - File management - πŸ“– V-Order Optimization - Performance optimization

Data Factory in Fabric

Data Pipelines - Copy data activity - Dataflow Gen2 for transformations - Pipeline orchestration - Schedule and triggers - Integration with Azure Data Factory - πŸ“– Data Factory Overview - Pipeline concepts - πŸ“– Copy Activity - Data ingestion - πŸ“– Dataflow Gen2 - Data transformation - πŸ“– Pipeline Activities - Activity reference - πŸ“– Pipeline Monitoring - Run monitoring - πŸ“– Pipeline Parameters - Parameterization

Notebooks and Spark

Fabric Notebooks - Python, Scala, R, SQL languages - Apache Spark engine - Built-in data visualization - Integration with lakehouses - Notebook scheduling - πŸ“– Notebooks Overview - Notebook guide - πŸ“– Spark Compute - Spark configurations - πŸ“– Notebook Source Control - Git integration - πŸ“– PySpark Reference - PySpark tutorial - πŸ“– Delta Lake with Spark - Delta operations

Data Warehouse

Fabric Warehouse - T-SQL analytics - Columnar storage - Separation of storage and compute - Automatic query optimization - Cross-database queries - πŸ“– Warehouse Overview - Warehouse concepts - πŸ“– Create Warehouse - Setup guide - πŸ“– Warehouse Tables - Table design - πŸ“– Warehouse Ingestion - Data loading - πŸ“– Query Warehouse - T-SQL queries - πŸ“– Warehouse Security - Row-level security

Semantic Models (Power BI Datasets)

Data Modeling - Import, DirectQuery, Composite modes - Star schema design - Relationships and cardinality - Calculated columns and measures - Hierarchies and display folders - πŸ“– Semantic Models Overview - Modeling concepts - πŸ“– DirectLake - Direct Lake mode - πŸ“– Data Modeling Best Practices - Star schema design - πŸ“– Relationships - Relationship types - πŸ“– Calculated Columns vs Measures - Calculations - πŸ“– Model Optimization - Performance tuning

DAX (Data Analysis Expressions)

DAX Functions and Patterns - Aggregation: SUM, AVERAGE, COUNT, MIN, MAX - Filter context: CALCULATE, FILTER, ALL, ALLEXCEPT - Time intelligence: TOTALYTD, SAMEPERIODLASTYEAR, DATEADD - Iterators: SUMX, AVERAGEX, COUNTX - Relationship navigation: RELATED, RELATEDTABLE - πŸ“– DAX Reference - Complete DAX reference - πŸ“– CALCULATE Function - Context transition - πŸ“– Time Intelligence - Date functions - πŸ“– DAX Best Practices - Performance patterns - πŸ“– Variables in DAX - VAR keyword

Real-Time Analytics (KQL Database)

Kusto Query Language (KQL) - Time-series analytics - Streaming data ingestion - Event-based data - Fast queries on large datasets - Integration with Event Hubs, IoT Hub - πŸ“– KQL Database Overview - KQL database - πŸ“– KQL Query Language - Query syntax - πŸ“– Data Ingestion - Eventstreams - πŸ“– KQL Queryset - Query management - πŸ“– Real-Time Dashboards - Dashboard creation

Power BI Integration

Reports and Dashboards - Report design and visualization - Paginated reports - Real-time streaming - Row-level security (RLS) - Apps and content distribution - πŸ“– Power BI in Fabric - Power BI integration - πŸ“– Create Reports - Report design - πŸ“– Visualizations - Visual types - πŸ“– Row-Level Security - RLS implementation - πŸ“– Paginated Reports - Report Builder - πŸ“– Power BI Apps - App deployment

Data Science in Fabric

Machine Learning - ML models with notebooks - MLflow integration - Model training and tracking - Model deployment - πŸ“– Data Science Overview - ML in Fabric - πŸ“– Train Models - Model training - πŸ“– MLflow - Experiment tracking - πŸ“– Model Registry - Model management

Security and Governance

Access Control - Workspace roles and permissions - Item permissions - Row-level security (RLS) - Object-level security (OLS) - Dynamic data masking - πŸ“– Fabric Security - Security overview - πŸ“– Workspace Permissions - Access control - πŸ“– Data Loss Prevention - Data protection - πŸ“– Endorsement - Content certification - πŸ“– Sensitivity Labels - Data classification

Monitoring and Optimization

Performance Monitoring - Monitoring hub - Capacity metrics app - Query performance analysis - Usage metrics - Diagnostic logs - πŸ“– Monitoring Hub - Centralized monitoring - πŸ“– Capacity Metrics - Capacity analytics - πŸ“– Query Insights - Query performance - πŸ“– Performance Analyzer - Report optimization

Data Integration Patterns

Common Patterns - Medallion architecture (Bronze, Silver, Gold) - Lambda architecture (batch + streaming) - Hub-and-spoke data distribution - Data mesh with domains - ELT over ETL - πŸ“– Medallion Architecture - Layered approach - πŸ“– Data Pipelines Best Practices - Design patterns

Migration to Fabric

Migration Strategies - Power BI workspace migration - Azure Synapse to Fabric - Azure Data Factory to Fabric Data Factory - On-premises data sources - πŸ“– Migrate to Fabric - Migration guide - πŸ“– Synapse Migration - Synapse to Fabric

Common Scenarios

Scenario 1: Modern Data Warehouse - Solution: Lakehouse β†’ Warehouse β†’ Semantic Model β†’ Power BI Reports

Scenario 2: Real-Time Analytics Dashboard - Solution: Eventstream β†’ KQL Database β†’ Real-Time Dashboard

Scenario 3: Self-Service BI - Solution: Lakehouse β†’ DirectLake Semantic Model β†’ Power BI Reports

Scenario 4: Data Engineering Pipeline - Solution: Source β†’ Data Pipeline β†’ Dataflow Gen2 β†’ Lakehouse (Bronze/Silver/Gold)

Scenario 5: ML Model Deployment - Solution: Lakehouse β†’ Notebook β†’ MLflow β†’ Model Registry β†’ Batch Scoring

Essential Commands

Python (PySpark) in Notebooks:

# Read from lakehouse
df = spark.read.format("delta").load("Tables/tablename")

# Write to lakehouse
df.write.format("delta").mode("overwrite").save("Tables/tablename")

# Optimize Delta table
spark.sql("OPTIMIZE tablename")

# V-Order write
df.write.format("delta").option("optimizeWrite", "true").save("Tables/tablename")

KQL Queries:

// Time-series aggregation
TableName
| where Timestamp > ago(1h)
| summarize Count=count() by bin(Timestamp, 5m)
| render timechart

// Top N analysis
TableName
| summarize Total=sum(Amount) by Category
| top 10 by Total desc

DAX Measures:

// Time intelligence
Sales YTD = TOTALYTD(SUM(Sales[Amount]), Dates[Date])

// Context modification
Sales All Regions = CALCULATE(SUM(Sales[Amount]), ALL(Region))

Exam Tips

Keywords: - "Unified analytics" β†’ Microsoft Fabric - "Single data lake" β†’ OneLake - "ACID transactions on data lake" β†’ Delta Lake format - "No data movement" β†’ Shortcuts - "Fast semantic model" β†’ DirectLake mode - "Real-time analytics" β†’ KQL Database, Eventstream - "T-SQL analytics" β†’ Warehouse - "Python/Spark transformations" β†’ Notebooks - "Low-code ETL" β†’ Dataflow Gen2

Focus Areas: - Understand when to use Lakehouse vs Warehouse - DirectLake mode and its benefits - OneLake shortcuts for data federation - Medallion architecture (Bronze, Silver, Gold layers) - DAX context (filter context vs row context) - KQL query syntax for time-series data - Workspace roles and permissions - Capacity management and optimization - Data modeling best practices

Study Strategy: - Hands-on: Create lakehouses, build data pipelines, design semantic models - Practice DAX and KQL queries - Understand DirectLake vs DirectQuery vs Import - Know the Fabric item types and their use cases - Memorize capacity SKUs and limits


Pro Tip: Microsoft Fabric is relatively new (GA November 2023). Focus on understanding the unified architecture, OneLake concepts, and how traditional Power BI/Synapse/Data Factory concepts map to Fabric!