DP-700 Sample Questions

DP-700 Sample Questions & Answers

Monitoring Fabric items, resolving errors and tuning performance carries slightly more weight than the rest, split between batch and streaming data ingestion and transformation, and configuring governance and lifecycle management around workspace settings.

Launch the full DP-700 simulator →

Showing 10 of 20 free samples.

  1. Question 1AdvancedSelect 2

    Implement and manage an analytics solution · Configure security and governance

    You are tasked with implementing security for a new Fabric Lakehouse. The requirements are as follows:

    1. Data engineers must have full control over the Lakehouse, including the ability to read and write to tables and underlying files.
    2. Business analysts must be able to query all data using the SQL endpoint but must be prevented from accessing the underlying files in OneLake.
    3. A specific group of junior analysts must only see data for the 'North America' region from the Sales table.

    Which combination of security mechanisms should you use? (Select TWO)

    Show answer & explanation

    Correct answers: A, C

    Workspace roles provide the broad level of access control. 'Admin' gives data engineers full control. 'Viewer' allows business analysts to read content, including querying the SQL endpoint, but prevents them from modifying items or accessing files directly, fulfilling the second requirement.

    Row-level security (RLS) is the correct mechanism for filtering rows based on user identity or role. You would create a security predicate (e.g., [Region] = 'North America') and apply it to a role that the junior analysts belong to, fulfilling the third requirement.

  2. Question 2Advanced

    Monitor and optimize an analytics solution · Optimize Spark performance

    You are troubleshooting a slow-running PySpark notebook in a Fabric workspace. The notebook reads a large Delta table, performs a series of transformations, and then joins it with a small dimension table. You observe from the Spark UI that one particular stage is taking an exceptionally long time and has a significant amount of data shuffle.

    The problematic code snippet is:
    large_df.join(small_df, 'user_id', 'inner')

    Which optimization technique should you apply to mitigate this performance bottleneck?

    Show answer & explanation

    Correct answer: C

    The scenario describes a classic large-table-to-small-table join, which is a perfect candidate for a broadcast hash join. By default, Spark might choose a shuffle-based join, causing massive data movement. Using from pyspark.sql.functions import broadcast and rewriting the join as large_df.join(broadcast(small_df), 'user_id', 'inner') instructs Spark to send the small DataFrame to every executor, avoiding the costly shuffle of the large DataFrame and dramatically improving performance.

  3. Question 3Intermediate

    Ingest and transform data · Process data by using KQL

    A manufacturing company uses Microsoft Fabric to monitor its production line. An Eventstream ingests sensor data, which is then written to a KQL database for real-time dashboarding. A critical requirement is to detect anomalies where the average temperature of a specific sensor (sensorId = 'SN01-T5') exceeds 150 degrees Celsius over a 5-minute period. You need to write a KQL query to create an alert based on this condition.

    Which KQL query correctly implements this logic?

    Show answer & explanation

    Correct answer: B

    This query correctly implements the logic. It first filters for the specific sensor. Then, summarize avg(temperature) by bin(timestamp, 5m) calculates the average temperature over 5-minute tumbling windows. Finally, where avg_temperature > 150 filters these aggregated results to find the windows where the anomaly condition is met.

  4. Question 4Beginner

    Ingest and transform data · Configure and use mirroring

    True or False: When you use the Mirroring feature in Microsoft Fabric to replicate an Azure SQL Database, the initial setup performs a full snapshot of the source data, and subsequent changes are then replicated in near real-time using change data capture (CDC), without requiring the configuration of a separate data pipeline.

    Show answer & explanation

    Correct answer: A

    This statement is true. The Fabric Mirroring feature is designed for seamless, low-latency replication. It automatically handles the initial snapshot of the source database tables and then uses the source's change data capture (CDC) mechanism to replicate ongoing inserts, updates, and deletes to OneLake in Delta format. This provides a continuously updated, queryable copy without the need to build and manage a manual ETL pipeline.

  5. Question 5Advanced

    Ingest and transform data · Design and implement data loading patterns

    Case Study: Global E-Commerce Platform

    Company Background
    GlobalCart is a multinational e-commerce company that operates in North America, Europe, and Asia. They are migrating their analytics platform to Microsoft Fabric to create a unified view of their sales, customer, and inventory data. Their primary goal is to empower regional business units with self-service analytics while maintaining centralized governance and data engineering standards.

    Existing Environment
    Data is currently stored in a mix of sources: an on-premises SQL Server 2019 database for inventory, an Azure SQL Database for sales transactions, and Parquet files in an ADLS Gen2 account for historical clickstream data. The company has a single Fabric F64 capacity located in 'Central US'. The data engineering team is proficient in SQL and PySpark.

    Requirements

    1. Architecture: Implement a medallion architecture using a central 'Corporate' workspace for Bronze and Silver layer Lakehouses. Each region (NA, EU, AS) must have its own separate workspace containing a Gold layer Lakehouse with data filtered specifically for that region.
    2. Data Ingestion: Data from the on-premises SQL Server must be ingested nightly with minimal impact on the source system. Sales data from Azure SQL should be replicated in near real-time.
    3. Transformation: Transformations from Bronze to Silver will be complex and require custom Python libraries. Transformations from Silver to the regional Gold layers will involve filtering and simple aggregations.
    4. Security: Data analysts in each region must only be able to access the Gold layer Lakehouse in their respective regional workspace. They should not have any access to the 'Corporate' workspace or other regions' workspaces.

    Problem Statement
    You need to design the data flow and transformation strategy to populate the regional Gold layer Lakehouses from the centralized Silver layer Lakehouse. The solution must be efficient and scalable. Which approach best meets the requirements?

    Show answer & explanation

    Correct answer: B

    This centralized approach aligns with the company's goal of maintaining central data engineering standards. A master pipeline in the 'Corporate' workspace can orchestrate the entire Silver-to-Gold process. Using notebooks for the transformation allows for the simple filtering and aggregation logic required. This design provides a single point of management and monitoring for the population of all Gold layers, ensuring consistency and leveraging the central team's expertise. It also avoids granting cross-workspace permissions to regional resources.

  6. Question 6Intermediate

    Implement and manage an analytics solution · Deploy items across workspaces by using deployment pipelines

    A data engineer has configured a Fabric deployment pipeline with Development, Test, and Production stages. After deploying changes from Test to Production, they notice that the connection details for a Lakehouse shortcut, which should point to a different storage account in Production, have not been updated. They need to ensure that connection details are automatically updated during deployment. What must be configured in the deployment pipeline?

    Show answer & explanation

    Correct answer: B

    Deployment rules are a feature of Fabric deployment pipelines specifically designed to handle differences in configuration between stages. You can create a rule for the Lakehouse shortcut that specifies the new connection details (e.g., the Production storage account URL) to be applied when deploying to the Production stage. This allows the same Fabric item to be used across stages while pointing to the correct environment-specific resources.

  7. Question 7Beginner

    Ingest and transform data · Configure and use Fabric shortcuts

    You are designing a solution to ingest data from an Amazon S3 bucket into a Fabric Lakehouse. The data consists of historical logs stored in Parquet format, totaling over 5 TB. The key requirements are to minimize data duplication and avoid cross-cloud data transfer costs for exploratory analysis, while still allowing data to be transformed and loaded into managed Delta tables for production reporting.

    Which Fabric feature should you use as the primary mechanism to access the S3 data?

    Show answer & explanation

    Correct answer: B

    OneLake shortcuts are the ideal solution for this scenario. A shortcut acts as a symbolic link to the data in the S3 bucket without copying it. This allows data engineers to query and explore the data in place using Spark or SQL, avoiding data duplication and egress costs for initial analysis. Later, a pipeline or notebook can read from the shortcut to transform and load the necessary data into managed Delta tables within the Lakehouse.

  8. Question 8Intermediate

    Monitor and optimize an analytics solution · Identify and resolve pipeline errors

    A data pipeline that loads data into a Fabric Warehouse fails intermittently. The pipeline monitoring view shows that the 'Copy Data' activity is failing with the error message: Execution Timeout Expired. The timeout period elapsed prior to completion of the operation or the server is not responding. The source is an on-premises SQL Server accessed via a data gateway.

    What is the most likely cause of this issue and the best first step to resolve it?

    Show answer & explanation

    Correct answer: B

    The error message Execution Timeout Expired points directly to a timeout on the command being executed against the source system. This is often due to a long-running source query, network latency over the data gateway, or load on the on-premises SQL Server. The most direct solution is to increase the 'Command timeout' property in the source dataset's advanced settings within the Copy activity to allow more time for the query to complete.

  9. Question 9Intermediate

    Monitor and optimize an analytics solution · Optimize Lakehouse table performance

    You are implementing a medallion architecture in a Fabric Lakehouse. You need to perform the VACUUM operation on your Delta tables to remove old, unreferenced files and manage storage costs. According to Microsoft's recommended best practices, how should you typically execute the VACUUM command?

    Show answer & explanation

    Correct answer: A

    The VACUUM command permanently deletes data. To prevent data loss from long-running queries or clock skew, there is a safety check that prevents you from vacuuming files newer than the retention period (default is 7 days or 168 hours). The recommended best practice is to schedule a notebook to run VACUUM periodically using the default retention period, ensuring a safe and consistent cleanup process.

  10. Question 10AdvancedSelect 2

    Ingest and transform data · Process data by using Spark structured streaming

    You are using Spark Structured Streaming in a Fabric notebook to process a continuous stream of JSON messages. The goal is to write the output to a Delta table while handling late-arriving data. You need to ensure that events arriving up to 1 hour late are included in the correct windowed aggregation. Which combination of Structured Streaming features should you use? (Select TWO)

    Show answer & explanation

    Correct answers: A, C

    Watermarking is Spark's mechanism for dealing with late-arriving data in stateful stream processing. By defining a watermark (e.g., withWatermark('eventTime', '1 hour')), you tell the engine how long to wait for late data before finalizing the state for a given window.

    The window() function is used to group data into time-based windows (e.g., tumbling, sliding). This is essential for performing aggregations over specific time intervals. Combining window() with withWatermark() allows for stateful, time-windowed aggregations that correctly handle late data.

Ready for the real thing?

The full DP-700 simulator has every exam-style question, timed mode, and instant scoring.

Go to the DP-700 simulator →