ADP Sample Questions

ADP Sample Questions & Answers

Free Associate Data Practitioner practice questions with worked answers and explanations. See how the ExamJungle simulator prepares you — then jump into the full test.

Launch the full ADP simulator →

Showing 10 of 20 free samples.

  1. Question 1Intermediate

    Data Management · Apply security measures and ensure compliance

    True or False: Using a customer-managed encryption key (CMEK) with Cloud Storage means that Google no longer holds any component of the encryption key, and all cryptographic operations happen outside of Google Cloud.

    Show answer & explanation

    Correct answer: B

    This statement is false. With CMEK, the customer manages the key within Google's Cloud Key Management Service (KMS). Google services (like Cloud Storage) still perform the cryptographic operations by requesting the KMS to use the key for encryption/decryption. The customer controls the key's lifecycle and access policies, but the key material resides within KMS. The scenario described (key held and used outside Google) applies to Customer-Supplied Encryption Keys (CSEK) or Cloud External Key Manager (EKM).

  2. Question 2Beginner

    Data Preparation and Ingestion · Data quality and cleaning

    A healthcare organization is building a data pipeline to process patient records. The raw data arrives as JSON files in a Cloud Storage bucket. The pipeline must de-identify sensitive information like patient names and social security numbers by applying masking transformations before loading the data into BigQuery for analysis. The solution must be a fully managed, graphical, low-code service to accelerate development. The pipeline should be designed according to the following flow:

    graph TD A[GCS Bucket: Raw JSON] --> B{De-identification Pipeline}; B --> C[BigQuery Table: Anonymized Data];

    Which service should be used to build the de-identification pipeline (B)?

    Show answer & explanation

    Correct answer: C

    Cloud Data Fusion is the ideal service for this requirement. It is a fully managed, cloud-native data integration service that provides a graphical interface and a broad library of pre-built transformations, including those for data masking and de-identification. This aligns perfectly with the 'low-code' and 'graphical' requirements. Dataflow would require custom coding, and Cloud Composer is an orchestrator, not a data transformation tool itself.

  3. Question 3Intermediate

    Data Analysis and Presentation · Define, train, evaluate, and use ML models

    You are building a regression model in BigQuery ML to predict housing prices. After training your model, you use the ML.EVALUATE function and get the following output: {'mean_absolute_error': 25000, 'r2_score': 0.85}. What do these metrics signify about your model's performance?

    Show answer & explanation

    Correct answer: C

    For a regression model, mean_absolute_error (MAE) indicates the average absolute difference between the predicted values and the actual values. An MAE of 25000 means the predictions are off by an average of $25,000. The r2_score (R-squared) represents the proportion of the variance in the dependent variable (housing price) that is predictable from the independent variables. An r2_score of 0.85 means that the model's features explain 85% of the variability in housing prices, which is generally considered a good fit.

  4. Question 4Beginner

    Data Pipeline Orchestration · Schedule, automate, and monitor basic data processing tasks

    An e-commerce company has a daily batch pipeline that updates product inventory in BigQuery. The pipeline is orchestrated by a Cloud Composer DAG. Recently, the DAG has been failing intermittently. You need to investigate the failures by reviewing the execution history, task logs, and the overall structure of the DAG. Which user interface should you use to perform this troubleshooting?

    Show answer & explanation

    Correct answer: C

    Cloud Composer is a managed Apache Airflow service. The primary interface for managing, monitoring, and troubleshooting DAGs (Directed Acyclic Graphs) is the Airflow web UI. This interface provides detailed views of DAG runs, task statuses, logs for individual tasks, trigger history, and the rendered code. While logs are also available in Cloud Logging and the environment can be managed in the Google Cloud Console, the Airflow UI is the specific tool designed for in-depth DAG troubleshooting.

  5. Question 5Intermediate

    Data Management · Configure access control and governance

    A new data analyst has joined your team and needs permissions to run queries on all tables within the production_analytics dataset in BigQuery. They also need to be able to create new tables in this dataset. However, they must not be able to delete the dataset or modify its permissions. Following the principle of least privilege, which single predefined IAM role should you grant to the analyst at the dataset level?

    Show answer & explanation

    Correct answer: B

    The BigQuery Data Editor role is the most appropriate choice. It grants permissions to read, query, create, update, and delete tables within a dataset. This meets the requirements of running queries and creating new tables. It does not, however, grant permissions to delete the dataset itself or change its IAM policies, which are part of the BigQuery Data Owner role. BigQuery Data Viewer is too restrictive as it only allows reading data, not creating tables.

  6. Question 6Intermediate

    Data Management · Configure lifecycle management

    A media company stores large video files in a Multi-Regional Cloud Storage bucket with Standard storage class. These files are accessed frequently for the first 30 days. After 30 days, access becomes rare, but the files must be kept for five years for compliance. After five years, they should be deleted. You need to implement a cost-effective lifecycle management policy. What should the lifecycle rule contain?

    Show answer & explanation

    Correct answer: B

    This configuration is the most cost-effective. After the initial 30 days of frequent access, the data becomes cold. Transitioning it directly to Archive storage, which is designed for long-term, infrequent access, will provide the lowest storage cost for the five-year retention period. A final action to delete the objects after 1825 days (365 * 5) fulfills the compliance requirement. Using Nearline or Coldline would be more expensive than Archive for this long-term storage use case.

  7. Question 7Intermediate

    Data Analysis and Presentation · Identify data trends, patterns, and insights by using BigQuery and Jupyter notebooks

    You are analyzing a dataset of customer transactions in a Colab Enterprise notebook. The data is stored in a BigQuery table. You need to identify the top 5 customers by total spending. The BigQuery table is named project.dataset.transactions and has columns customer_id and purchase_amount. You have already authenticated and initialized the BigQuery client in your notebook. Which Python code snippet using the BigQuery client library should you use?

    Show answer & explanation

    Correct answer: A

    This is the correct and standard way to execute a SQL query against BigQuery from a Python environment using the official client library. The SQL query correctly calculates the sum of purchase_amount for each customer_id, orders the results in descending order to find the top spenders, and limits the output to the top 5. The client.query(sql).to_dataframe() method executes the query and conveniently converts the results into a Pandas DataFrame for further analysis in the notebook.

  8. Question 8Intermediate

    Data Pipeline Orchestration · Identify use cases for event-driven data ingestion from Pub/Sub to BigQuery

    A startup is building an event-driven data processing system. When a new JSON file is uploaded to a Cloud Storage bucket, a lightweight data validation check needs to be performed. If the file is valid, a message should be sent to a Pub/Sub topic for downstream processing. The solution must be serverless, cost-effective for infrequent uploads, and require minimal infrastructure management. Which combination of services should be used to build this system?

    sequenceDiagram participant User participant GCS as Cloud Storage participant Trigger participant Processor participant PubSub as Pub/Sub User->>GCS: Upload JSON file GCS->>Trigger: Event: object.finalize Trigger->>Processor: Invoke with event data Processor->>Processor: Validate JSON Processor->>PubSub: Publish message
    Show answer & explanation

    Correct answer: B

    This is the classic serverless pattern for lightweight, event-driven tasks on Google Cloud. A Cloud Function can be directly triggered by a Cloud Storage event (like a file upload). The function can execute the validation logic and then publish to Pub/Sub. This solution is fully serverless, scales to zero (making it cost-effective for infrequent events), and requires no infrastructure management. Dataflow is better suited for large-scale stream or batch processing, not single-file validation. Cloud Run is also a good serverless option but Cloud Functions are often simpler for direct event handling like this.

  9. Question 9Advanced

    Data Preparation and Ingestion · Data quality and cleaning

    During a data quality audit, you discover that a critical customers table in BigQuery contains duplicate rows based on the customer_id column. You need to create a new table, customers_deduped, that contains only the unique customer records. For duplicates, you must keep the record that was most recently updated, based on the last_modified_ts timestamp column. Which SQL query should you use?

    Show answer & explanation

    Correct answer: B

    This query uses a window function ROW_NUMBER() to rank rows for each customer_id based on the last_modified_ts in descending order. This assigns a rank of 1 to the most recent record for each customer. The QUALIFY clause, specific to BigQuery and some other data warehouses, filters the results of the window function, keeping only the rows where the rank is 1. This is an efficient and standard pattern for deduplication in BigQuery. The other options use incorrect SQL syntax or logic for this task.

  10. Question 10IntermediateSelect 2

    Data Management · Identify high availability and disaster recovery strategies

    A gaming company is designing a high-availability and disaster recovery (HA/DR) strategy for its player profile data, which is stored in Cloud SQL for PostgreSQL. The primary business requirements are to ensure data durability against regional failures and to provide a read-only endpoint for analytics that does not impact the primary instance's performance. Which Cloud SQL features should be configured to meet these requirements? (Select TWO)

    Show answer & explanation

    Correct answers: A, B

Ready for the real thing?

The full ADP simulator has every exam-style question, timed mode, and instant scoring.

Go to the ADP simulator →