GEN-AI-ENG Sample Questions & Answers
Creating and selecting tools alongside prompt optimization is the single biggest topic, alongside assembling and deploying RAG chains, evaluating and monitoring models, model task selection, document chunking for retrieval, and governance guardrails.
Launch the full GEN-AI-ENG simulator →Showing 10 of 20 free samples.
- Question 1Beginner
Data Preparation · Define operations and sequence to write given chunked text into Delta Lake tables in Unity Catalog
A data engineer needs to load processed text chunks into a Delta table for a RAG application. The data is currently in a Spark DataFrame named
chunks_dfwith columnsdoc_id,chunk_text, andchunk_sequence. The target table,workspace.default.document_chunks, must be created if it does not exist and overwritten if it does. Which command correctly performs this operation?Show answer & explanation
Correct answer: A
This command correctly uses the Spark DataFrameWriter API.
format("delta")specifies the storage format.mode("overwrite")ensures that if the table already exists, its contents are replaced, and if it doesn't, it will be created.saveAsTable()is the action that writes the data to the specified table name in Unity Catalog. - Question 2Intermediate
Design Applications · Design a prompt that elicits a specifically formatted response
A developer is creating a prompt for a marketing campaign slogan generator. The business requires the output to be a JSON object containing three distinct slogan options, each with a specific 'style' (e.g., 'playful', 'professional', 'bold'). Which prompt design is most likely to elicit the desired structured response consistently?
Show answer & explanation
Correct answer: C
To get consistent, structured output like JSON, the most reliable method is few-shot (or in this case, one-shot) prompting. By providing a concrete example of the desired output format, you are giving the model a clear template to follow. This significantly reduces the ambiguity and variability in its response, making it much more likely to generate a correctly formatted JSON object with the specified keys and structure every time.
- Question 3Intermediate
Assembling and Deploying Applications · Identify batch inference workloads and apply ai_query() appropriately
A financial firm is deploying a sentiment analysis model on earnings call transcripts. To manage costs, they want to use the
ai_query()function for batch processing directly within their Databricks SQL workflow. The model is served at an endpoint namedsentiment_analyzer. The transcripts are in a tabletranscriptswith a columntranscript_text. What is the correct SQL syntax to invoke the model?Show answer & explanation
Correct answer: A
The
ai_query()function is the standard way to invoke a model serving endpoint from within Databricks SQL for batch inference. The correct syntax requires the endpoint name as the first argument and the input data (in this case, thetranscript_textcolumn) as the second argument. This allows for seamless, scalable inference on large datasets directly within a SQL query. - Question 4Beginner
Evaluation and Monitoring · Use inference logging to assess deployed RAG application performance
True or False: When enabling Inference Tables on a Databricks Model Serving endpoint, both the requests and the responses are automatically captured and stored in a Delta table in the user's Unity Catalog schema without requiring any code changes to the client application.
Show answer & explanation
Correct answer: A
This statement is true. Inference Tables are a managed feature of Databricks Model Serving. When enabled on an endpoint, Databricks automatically captures the payload of incoming requests and outgoing responses and logs them to a managed Delta table. This process is transparent to the client application and requires no modification to the code making the API calls.
- Question 5Intermediate
Assembling and Deploying Applications · Code a chain using a pyfunc model with pre- and post-processing
A developer is building a custom
pyfuncmodel for a RAG chain. The model needs to perform a pre-processing step to extract keywords from the user's query before sending it to the retriever. The keyword extraction logic is contained in a helper function within a separate Python file (utils.py). How should the developer package the model with MLflow to ensure theutils.pyfile is available at inference time?Show answer & explanation
Correct answer: C
The
code_pathparameter inmlflow.pyfunc.log_modelis specifically designed for this purpose. It takes a list of file paths or directories. MLflow will automatically package these files with the model and add them to the Python path when the model is loaded for inference. This ensures that any custom modules or helper functions are available to the model's code. - Question 6Intermediate
Governance · Recommend an alternative for problematic text mitigation in a data source feeding a GenAI applicatioN
A hospital is deploying a GenAI application to help doctors summarize patient visit notes. To comply with HIPAA, the application must prevent any possibility of leaking patient data into model training logs or to unauthorized users. Which of the following is the most critical guardrail to implement?
Show answer & explanation
Correct answer: B
Relying on a prompt to protect sensitive data is insufficient and unreliable for compliance. The most robust solution is a programmatic guardrail that actively detects and removes or masks PII from the input before it reaches the LLM. This prevents the sensitive data from ever being processed or logged by the model provider. A corresponding check on the output adds another layer of security.
- Question 7Intermediate
Design Applications · Define and order tools that gather knowledge or take actions for multi-stage reasoning
An engineer is building an agent-based system to help users plan a vacation. The agent needs to perform a sequence of tasks: ask the user for their destination and dates, find available flights, find hotels, and then present a combined itinerary. Which components are essential for defining this multi-stage reasoning process in a framework like LangChain?
Show answer & explanation
Correct answer: B
Agent-based systems function by breaking a complex task into smaller, manageable steps. This requires: 1) A set of discrete
Toolsthat can perform specific actions (like calling an API). 2) AnLLMthat acts as the 'brain' to decide which tool to use next based on the current state. 3) AnAgent Executorthat manages the loop of running the chosen tool, observing the output, and feeding it back to the LLM until the final goal is reached. - Question 8Beginner
Application Development · Select a embedding model context length based on source documents, expected queries, and optimization strategy
A team is building a RAG application and needs to select an embedding model. The source documents are short, averaging about 200 tokens. The primary constraints are minimizing inference latency and operational cost. Which embedding model profile would be the best choice?
Show answer & explanation
Correct answer: B
For minimizing latency and cost, a smaller model is always better. A context length of 512 is more than sufficient for documents averaging 200 tokens, so a larger context window would be wasteful. A smaller embedding dimension (384) results in a smaller model size, faster inference, and lower storage costs in the vector database. Larger models with more dimensions should only be chosen when maximum quality is the primary goal.
- Question 9Intermediate
Evaluation and Monitoring · Use Databricks features to control LLM costs for RAG applications
An operations team is tasked with controlling the costs of a suite of GenAI applications. One application, used for summarizing daily news, is driving a significant portion of the token costs. The summaries are useful but do not need to be exhaustive. Which feature within Databricks is specifically designed to help manage and cap the costs associated with Foundation Model API usage?
Show answer & explanation
Correct answer: C
Databricks Model Serving provides built-in cost management features for endpoints that use Foundation Model APIs. Administrators can configure specific rate limits (e.g., tokens per minute) and hard cost budgets (e.g., dollars per hour/day/month). Once a budget is exhausted, the endpoint will stop serving requests, providing a reliable mechanism to prevent cost overruns.
- Question 10Intermediate
Assembling and Deploying Applications · Create and query a Vector Search index
A developer is creating a Vector Search index for a RAG application. The application requires near real-time updates as new documents are added to the source Delta table. Which Vector Search indexing mode should be selected to meet this requirement?
Show answer & explanation
Correct answer: D
The Delta Sync indexing mode is specifically designed for this use case. It automatically and continuously monitors a source Delta table for changes (appends, updates, deletes) and incrementally updates the Vector Search index to reflect those changes. This provides a low-latency, near real-time synchronization between the source data and the searchable index, which is ideal for applications requiring up-to-date information.
Ready for the real thing?
The full GEN-AI-ENG simulator has every exam-style question, timed mode, and instant scoring.