GCP-ADP Sample Questions & Answers
Getting data ready and loaded into Google Cloud storage carries the biggest weight, alongside spotting trends with BigQuery and Jupyter, visualizing them in Looker, training basic ML models, building pipelines, and governance and recovery planning.
Launch the full GCP-ADP simulator →Showing 6 of 12 free samples.
- Question 1Beginner
Data Preparation and Ingestion · Prepare and process data
A marketing team needs to analyze customer sentiment from social media posts. The data is unstructured text stored in Cloud Storage. You need to process this data, extract sentiment scores, and store the structured results in BigQuery for analysis. Which tool provides a visual interface to build this pipeline without writing extensive code?
Show answer & explanation
Correct answer: C
Cloud Data Fusion is a fully managed, cloud-native data integration service that provides a visual point-and-click interface (GUI) for building ETL/ELT pipelines. It is ideal for users who want to build pipelines without writing code (code-free). It can leverage plugins to process text and write to BigQuery.
- Question 2Intermediate
Data Management · Configure lifecycle management
You are managing a data lake in Cloud Storage containing sensitive customer information. You need to ensure that the data is stored cost-effectively. The data is accessed frequently during the first 30 days, rarely for the next 60 days, and almost never after 90 days, but must be retained for 3 years for compliance. Which Object Lifecycle Management configuration should you apply?
Show answer & explanation
Correct answer: B
This matches the access patterns described: Standard for frequent access ( Nearline -> Coldline is the most logical step-down cost optimization for the described timeline.
- Question 3Intermediate
Data Analysis and Presentation · Identify trends and insights
You are analyzing a large dataset in BigQuery to identify sales trends. You need to calculate the moving average of sales over the last 7 days for each product. Which SQL construct should you use?
Show answer & explanation
Correct answer: B
Window functions allow calculations across a set of table rows that are somehow related to the current row. To calculate a moving average, you use the
OVERclause to define the window.PARTITION BYseparates the calculation by product,ORDER BYensures chronological order, andROWS BETWEEN 6 PRECEDING AND CURRENT ROWdefines the 7-day window (current day + previous 6 days). - Question 4Beginner
Data Analysis and Presentation · Visualize data and create dashboards
A data analyst needs to create an ad-hoc report to visualize sales performance for a specific region. The report is for a one-time presentation and does not require complex data modeling or version control. Which tool is most appropriate for this task?
Show answer & explanation
Correct answer: B
Looker Studio (formerly Google Data Studio) is a free, self-service business intelligence tool that is ideal for creating ad-hoc reports and dashboards quickly. It connects easily to data sources like BigQuery and Sheets without requiring the semantic modeling layer (LookML) or infrastructure setup of enterprise Looker.
- Question 5Advanced
Data Management · Configure access control and governance
You are setting up a new BigQuery dataset that will contain highly sensitive PII data. You need to ensure that only a specific group of Data Scientists can run queries against this dataset, while another group of Auditors can only view the metadata (table schemas) but not the data itself. Which IAM roles should you assign to each group for this dataset?
Show answer & explanation
Correct answer: B
To run queries, a user needs both permission to read the data (
dataViewer) and permission to run query jobs (jobUserproject-level or specific). For the Auditors who need to see schema but not data,roles/bigquery.metadataVieweris the principle of least privilege, asdataViewerwould allow them to see the table contents. - Question 6Advanced
Data Pipeline Orchestration · Schedule, automate, and monitor
You are designing a data pipeline that uses Cloud Dataflow to process streaming data from Pub/Sub and write it to BigQuery. You notice that the pipeline is experiencing high latency and the system lag is increasing. You inspect the logs and see many 'Hot Key' warnings. What is the most likely cause?
Show answer & explanation
Correct answer: C
A 'Hot Key' issue occurs when a specific key in your data (e.g., a specific user ID or category) appears much more frequently than others. In a parallel processing system like Dataflow, data is partitioned by key. If one key has a massive amount of data, the worker assigned to that key becomes a bottleneck, processing far more data than others, leading to high latency and lag. This is a classic data skew problem.
Ready for the real thing?
The full GCP-ADP simulator has every exam-style question, timed mode, and instant scoring.