Key details for this exam, checked against the published exam outline
Each question shows the correct answer and an explanation of why it is right
You need to configure the telemetry pipeline to support the planned changes for pipeline orchestration and address the resiliency issues.
What should you do?
Lakeflow Jobs provides native orchestration for multi-task Databricks workflows. Separate ingestion, cleansing, and curation tasks can be connected through explicit dependencies, ensuring that each stage starts only after its required upstream work succeeds. Each task can also have independent retry, notification, timeout, and compute settings, directly addressing the pipeline's resiliency requirements. Azure Data Factory could orchestrate notebooks, but it introduces another service when Lakeflow Jobs already provides the required functionality. A single notebook makes failures harder to isolate and can force successful stages to be rerun. Independently scheduled jobs rely on timing assumptions rather than actual task completion and can fail when an upstream stage runs longer than expected. Explicit Lakeflow Jobs dependencies provide reliable execution order and centralized monitoring. Microsoft Learn
You have an Azure Databricks workspace that is enabled for Unity Catalog.
You have 500 GB of sales data stored as multiple CSV files in cloud storage.
You plan to load the data into a Delta table.
You need to ingest the bulk data by using a solution that meets the following requirements:
* Minimize how long it takes to implement the solution.
* Minimize the amount of custom code required.
What should you use?
COPY INTO provides a concise SQL-based mechanism for loading files from cloud storage directly into a Delta table. It requires substantially less custom code than constructing a Spark ingestion application and is suitable for a straightforward bulk load of multiple CSV files. COPY INTO is also retryable and idempotent: it tracks files already loaded into the target table and skips them during later executions, helping prevent accidental duplication. Auto Loader is optimized primarily for incremental and continuously arriving files and normally requires a streaming or triggered pipeline. Apache Spark read APIs require additional code for reading, transforming, tracking, and writing the files. Manually uploading 500 GB would be inefficient and operationally unsuitable. Therefore, COPY INTO best satisfies both implementation-speed and minimal-code requirements. Microsoft Learn
You have an Azure Databricks workspace that is enabled for Unity Catalog.
You have a Lakeflow Spark Declarative Pipelines (SDP) pipeline that writes numerical data to a table named Table1 by using a data quality validation rule named rule1.
You need to modify rule1 to meet the following requirements:
Ensure that amount is always greater than 0.
Prevent an update to Table1 from being committed when data that violates rule1 is detected.
Which statement should you execute?
The correct answer is C --- @dlt.expect_or_fail.
Lakeflow Spark Declarative Pipelines (SDP) offers three expectation decorators, each with a different violation response:
@dlt.expect --- logs the violation as a metric but writes all records, including bad ones, to the table. Suitable for monitoring only.
@dlt.expect_or_drop --- drops violating records and continues the pipeline. The table receives only clean rows, but the pipeline update commits successfully.
@dlt.expect_or_fail --- fails the entire pipeline update when a violation is detected. The table update is never committed. This is the correct choice when data integrity is non-negotiable: 'Prevent an update to Table1 from being committed when data that violates rule1 is detected.'
@dlt.expect_all_or_drop takes a dictionary of rules and drops violating rows but still commits --- it doesn't halt the pipeline.
You have an Azure Databricks workspace that is enabled for Unity Catalog and contains a Delta table named Orders.
You load the Orders table into an Apache Spark DataFrame named df.
You need to create a DataFrame that excludes rows where the order amount is null.
Solution: You run the following expression.
df.dropna(subset=["order_amount"])
Does this meet the goal?
The correct answer is A --- Yes.
df.dropna(subset=['order_amount']) is the idiomatic PySpark way to remove rows where a specific column contains a null. It inspects only the columns listed in subset and drops any row where those columns are null. The resulting DataFrame contains only rows where order_amount is not null --- exactly what the requirement asks for.
The subset parameter is important: without it, dropna() would drop rows where ANY column is null, which could incorrectly exclude rows that have nulls in other columns but a valid order_amount. By specifying subset=['order_amount'], the filter is applied precisely and only to the column in question.
This method is semantically equivalent to df.filter(df.order_amount.isNotNull()) and to the SQL clause WHERE order_amount IS NOT NULL. Both are correct --- dropna with a subset is arguably the more readable Pythonic approach.
You have an Azure Databricks account that contains workspaces enabled for Unity Catalog.
You need to implement audit logging to meet the following requirements:
* Capture audit logs for all the workspaces in the account.
* Retain the audit logs for 90 days.
* Minimize storage and ingestion costs.
The logs will be reviewed only during security investigations and will NOT be queried regularly.
To where should you send the audit logs?
An Azure Storage account provides durable, comparatively low-cost retention for diagnostic and audit logs that are accessed infrequently. A lifecycle or retention policy can preserve the logs for 90 days and then remove them automatically. This matches an investigation-only access pattern without paying the ingestion and indexing charges associated with Log Analytics. Azure Monitor metrics stores numerical monitoring measurements, not the complete audit-event records required here. Azure Event Hubs is a streaming transport intended to forward events to consumers and is not the final long-term retention destination. Log Analytics is appropriate when teams need frequent querying, dashboards, and alerting, but those capabilities introduce unnecessary cost for logs reviewed only during occasional investigations. Storage therefore best satisfies centralized retention and cost requirements.
You have an Azure Databricks workspace that uses Unity Catalog.
You have a Lakeflow Spark Declarative Pipelines (SDP) pipeline that ingests data into a managed Delta table named Table1. Table1 is used for analytics.
New columns are added to the source data, causing pipeline failures during writes to Table1.
You need to prevent the pipeline failures. The solution must ensure that schema changes are detected and handled.
What should you do?
Schema evolution allows the target Delta table to incorporate compatible new source columns instead of failing when the incoming schema changes. This is the appropriate response to additive schema drift and avoids manually rebuilding tables whenever the source evolves. Creating a separate table for every schema version would fragment the dataset and increase operational effort. Disabling schema enforcement removes valuable protection against incompatible or corrupt data rather than handling legitimate evolution safely. Row filters operate on records and cannot remove an unexpected column from the incoming schema. With schema evolution enabled, the pipeline can detect new fields, update the target schema, and continue processing while retaining Delta Lake's transactional guarantees. The solution therefore supports changing source data without sacrificing the managed-table architecture.
Exam domains verified against: Official Microsoft DP-750 exam guide, last checked September 2026.
Choose appropriate compute types including job compute, serverless, warehouse, classic compute, and shared compute. Configure compute performance and features including CPU, node count, autoscaling, Photon acceleration, and runtime versions. Create and organize Unity Catalog objects through naming conventions, catalogs, schemas, volumes, tables, views, and materialized views. Implement foreign catalogs and DDL operations.
Sample question from this domain above: Q4
Grant privileges to principals for Unity Catalog objects and implement table, column, and row-level access control. Access Azure Key Vault secrets and authenticate using service principals and managed identities. Create descriptions for data discovery, configure attribute-based access control using tags and policies, apply row filters and column masks, manage data retention, set up lineage tracking, and configure audit logging.
Sample question from this domain above: Q2
Design data modeling including ingestion logic, extraction types, and file formats. Choose ingestion tools like Lakeflow Connect, notebooks, and Azure Data Factory, plus loading methods and table formats like Parquet, Delta, CSV, JSON, or Iceberg. Design partitioning schemes, SCD types, and clustering strategies. Ingest data using Lakeflow Connect, notebooks, SQL methods, CDC feeds, and Spark Structured Streaming from Event Hubs. Cleanse and transform data through profiling, type selection, handling duplicates and nulls, and using joins and aggregations. Implement data quality validation including nullability, cardinality, range checks, and schema enforcement.
Sample question from this domain above: Q1
Design pipeline order of operations and choose between notebooks and Lakeflow Spark Declarative Pipelines. Create jobs with setup, configuration, triggers, schedules, and alerts. Implement error handling and automatic restarts. Apply Git version control with branching and pull requests. Implement testing strategies including unit, integration, end-to-end, and UAT tests. Configure and deploy Databricks Asset Bundles using CLI and REST APIs. Monitor cluster consumption, troubleshoot Lakeflow Jobs and Spark jobs, investigate caching and shuffle issues using Spark UI, optimize Delta tables with OPTIMIZE and VACUUM, and configure logging and alerts in Azure Monitor.
Common questions about the exam itself