C1000-185 Sample Questions & Answers
Fine-tuning, including hard versus soft prompts and cost reduction, dominates the exam, ahead of core generative AI and LLM capabilities, zero- and few-shot prompt engineering, retrieval-augmented generation, deployment, and orchestrating APIs.
Launch the full C1000-185 simulator →Free C1000-185 Sample Questions with Answers
Real questions from the IBM watsonx Generative AI Engineer v1 - Associate practice test — answers and explanations included. Showing 20 of 40 free samples.
- Question 1Intermediate
Analyze and Design a Generative AI Solution · Understand security risks associated with LLMs, prompt engineering, prompt, and data
A public-facing chatbot built on a foundation model is found to be vulnerable to indirect prompt injection. An attacker is able to embed malicious instructions in a document that the chatbot later retrieves and processes, causing it to exfiltrate user data. Which of the following is the most effective mitigation strategy against this type of attack?
Show answer & explanation
Correct answer: D
Indirect prompt injection occurs when the malicious instruction comes from a data source (like a retrieved document) rather than the direct user input. Therefore, input validation on the user's prompt is ineffective. The best defense is to architect the prompt template to create a clear separation of concerns. The system prompt should explicitly instruct the LLM that the retrieved content is for informational purposes only and that any instructions within it should be ignored. This creates a logical barrier, making it much harder for the model to be manipulated by the content it processes.
- Question 2Intermediate
Prompt Engineering · Generate prompt templates
When creating a reusable prompt template in the watsonx.ai Prompt Lab for generating product descriptions, a developer wants to dynamically insert the product name and its key features. The template looks like this:
Generate a compelling marketing description for a new product named {{product_name}}. Highlight the following key features: {{features}}.To use this template via the API, the product name and features would be passed as values for the _____.
Show answer & explanation
Correct answer: B
The
{{product_name}}and{{features}}placeholders in the template are defined as prompt variables. When making an API call or using the template, the developer provides the actual values for these variables, which are then substituted into the template to form the final prompt sent to the model. - Question 3Advanced
Fine-Tuning · Prepare the dataset for training
A healthcare provider is developing a generative AI application to summarize clinician's notes into a patient-friendly format. The application must adhere to strict data privacy regulations (e.g., HIPAA) and must not send any sensitive patient data to external, third-party model providers. The provider has a large corpus of anonymized clinician notes and corresponding patient-friendly summaries to use for training.
The IT department has provisioned a secure, on-premises environment with powerful GPUs. The goal is to create a highly specialized model that excels at this specific summarization task and can be hosted entirely within their own infrastructure. The development team is evaluating different approaches on the watsonx platform.
Given the strict privacy constraints and the availability of a high-quality, task-specific dataset, what is the most appropriate strategy?
Show answer & explanation
Correct answer: C
This strategy directly addresses all key requirements. Fine-tuning (either full or PEFT) is the best method for creating a model that is highly specialized for a specific task when a quality dataset is available. It will yield superior performance compared to prompting alone. Most importantly, the resulting custom model is an asset that can be deployed entirely on-premises, satisfying the strict data privacy and security constraints by ensuring no sensitive data ever leaves their controlled environment.
- Question 4IntermediateSelect 3
Retrieval-Augmented Generation (RAG) · Develop using libraries
When designing a RAG pipeline using LangChain to work with watsonx.ai, which THREE of the following components are essential for the retrieval and generation process? (Select THREE)
flowchart TD A[Load Documents] --> B{Split into Chunks} B --> C[Generate Embeddings] C --> D[(Store in Vector DB)] E[User Query] --> F[Generate Query Embedding] F --> G{Search Vector DB} G --> H[Retrieve Relevant Chunks] H & E --> I{Construct Prompt} I --> J[Invoke LLM] J --> K[Generated Response]Show answer & explanation
Correct answers: A, C, D
A Document Loader is the first step in the RAG pipeline, responsible for ingesting data from various sources (PDFs, websites, databases) into a format LangChain can process.
A Vector Store (like Chroma or Milvus) stores the document embeddings. A Retriever is the LangChain component that interfaces with the Vector Store to find and return the most relevant document chunks based on the user's query.
This is the core generative component. LangChain uses wrappers for various LLMs (including those on watsonx.ai) to provide a standardized interface for sending the combined prompt (user query + retrieved context) and receiving the final generated answer.
- Question 5Intermediate
Deployment · Deploy a custom model
After successfully fine-tuning a foundation model for a specific task, an engineer needs to make it available for other developers in the organization to use via a REST API. What is the standard process for deploying this custom model as an endpoint within watsonx.ai?
Show answer & explanation
Correct answer: B
This describes the standard MLOps workflow in watsonx.ai and IBM Cloud Pak for Data. Models and other assets are developed within a project. To make them operational, they are promoted to a dedicated deployment space. From there, an online deployment can be created, which provisions the necessary resources and exposes the model via a stable, scalable REST API endpoint.
- Question 6Advanced
Analyze and Design a Generative AI Solution · Understand how to choose the appropriate model for a use case
A startup is developing a code generation assistant for a niche programming language. They have a limited budget and GPU capacity. Their primary requirements are low-latency suggestions and the ability to run the model on developer machines with moderate resources. They are choosing between different sizes of the IBM Granite Code models. Which model would be the most appropriate choice given these constraints?
Show answer & explanation
Correct answer: C
For this use case, the constraints of budget, GPU capacity, low latency, and running on developer machines are paramount. The smallest instruction-tuned code model is the optimal choice. Smaller models are significantly cheaper to run, have lower latency, and require less memory/GPU, making them suitable for local deployment. While the largest model might offer slightly higher quality, the operational costs and resource requirements would be prohibitive for the startup's constraints. An instruction-tuned model is also crucial for a chat/assistant-like application.
- Question 7Intermediate
Prompt Engineering · Determine the best model parameters for each GenAI prompt
In prompt engineering, both Top-P (nucleus) sampling and Top-K sampling are used to control the randomness of a model's output by limiting the pool of candidate tokens. What is the key difference in how they operate?
Show answer & explanation
Correct answer: C
This is the core difference. Top-K sampling considers a static pool size (e.g., the top 50 most likely tokens). Top-P sampling is dynamic; it considers the smallest set of tokens whose cumulative probability is greater than or equal to the value 'p'. If the model is very certain about the next token, this set might be very small (e.g., 2-3 tokens). If the model is uncertain, the set could be much larger. This makes Top-P generally more adaptive and often preferred over Top-K.
- Question 8Advanced
Fine-Tuning · LoRA
When using LoRA for parameter-efficient fine-tuning, what is the primary trade-off an engineer must consider when selecting the value for the rank (r)?
Show answer & explanation
Correct answer: C
The rank (r) directly controls the size of the low-rank adaptation matrices (A and B). A higher 'r' means larger matrices, which translates to more trainable parameters. This allows the model to learn more complex adaptations (higher expressiveness), but it also increases the computational and memory requirements during training. Furthermore, a rank that is too high relative to the size and complexity of the fine-tuning dataset can lead to overfitting, where the model memorizes the training data instead of generalizing.
- Question 9Intermediate
Retrieval-Augmented Generation (RAG) · Generate vector embeddings utilizing models
A multinational corporation is building a RAG system to serve employees in both North America and Japan. The system needs to process and retrieve information from technical manuals written in both English and Japanese. Which embedding model available in the watsonx.ai catalog would be the most suitable choice for this task?
Show answer & explanation
Correct answer: D
For a RAG system that must handle multiple languages, a dedicated multilingual embedding model is essential. Models like
paraphrase-multilingual-mpnet-base-v2are trained on many languages and map semantically similar sentences to nearby points in the vector space, regardless of the source language. This allows a single vector store to handle documents in both English and Japanese, and enables cross-lingual retrieval (e.g., asking a question in English and retrieving a relevant Japanese document). Using separate models would be complex and would not support cross-lingual search. - Question 10Beginner
Deployment · Deploy AI Assets
True or False: In watsonx.ai, only fine-tuned custom models can be promoted to a deployment space and deployed as AI assets.
Show answer & explanation
Correct answer: B
False. An 'AI Asset' in watsonx.ai is a broad term. Besides custom models, other assets such as saved prompt templates from the Prompt Lab, Python functions, and data assets can also be promoted to a deployment space and deployed for operational use.
- Question 11Advanced
Analyze and Design a Generative AI Solution · Articulate the optimal model architecture based on a use case
A financial services firm is designing a generative AI solution to summarize quarterly earnings reports. The reports are highly confidential, and the firm has a strict policy against data leaving their private cloud environment. They also require the model to cite specific numbers from the reports to verify accuracy. Which architectural pattern and deployment strategy best fits these requirements?
Show answer & explanation
Correct answer: B
The requirements specify strict data privacy (on-premises/private cloud), confidentiality, and the need for factual citations. RAG is the optimal pattern for grounding answers in specific documents (reducing hallucinations and enabling citations). Using an on-premises deployment of watsonx.ai satisfies the data sovereignty requirement, and IBM Granite models are enterprise-grade foundation models suitable for business summarization tasks.
- Question 12Beginner
Analyze and Design a Generative AI Solution · Understand how to choose the appropriate model for a use case
You are analyzing a use case for a coding assistant tool within a software development company. The developers primarily use Python and Java. You need to select a foundation model available in watsonx.ai that is specifically optimized for code generation tasks. Which model family should you prioritize for evaluation?
Show answer & explanation
Correct answer: B
The IBM Granite model family includes specific variants trained on code datasets (granite-code). The
granite-20b-codemodel is specifically designed and optimized for code generation, translation, and explanation tasks in languages like Python and Java. FLAN-T5 is an instruction-tuned text model, and while Llama and Elyza have capabilities, Granite Code is the native IBM enterprise option for this specific domain. - Question 13IntermediateSelect 2
Analyze and Design a Generative AI Solution · Understand the limitations of GenAI/LLMs
During the design phase of a customer service chatbot, you identify a risk that the Large Language Model (LLM) might confidently generate incorrect instructions for troubleshooting hardware devices. This phenomenon is known as hallucination. Which two design strategies are most effective at mitigating this specific risk? (Select TWO)
Show answer & explanation
Correct answers: B, C
RAG provides the model with factual context that it must use to generate the answer, significantly reducing hallucinations. Lowering temperature (decoding parameters) makes the model more deterministic and less likely to 'invent' creative but incorrect facts. Increasing temperature would actually increase hallucination risk.
RAG provides the model with factual context that it must use to generate the answer, significantly reducing hallucinations. Lowering temperature (decoding parameters) makes the model more deterministic and less likely to 'invent' creative but incorrect facts. Increasing temperature would actually increase hallucination risk.
- Question 14Beginner
Prompt Engineering · Determine the best model parameters for each GenAI prompt
A developer is using the watsonx.ai Prompt Lab to create a prompt for summarizing email threads. The model often stops generating text in the middle of a sentence. Which parameter should the developer adjust to ensure the model completes the summary while preventing it from rambling indefinitely?
Show answer & explanation
Correct answer: C
If a model stops mid-sentence, it likely hit the 'Max Tokens' limit. Increasing this allows for longer generation. To prevent indefinite rambling when the limit is raised, 'Stop Sequences' (e.g., a period, newline, or specific token like 'END') should be configured to signal the model to cease generation at a logical point.
- Question 15Beginner
Prompt Engineering · Describe the benefits of using prompt variables
You are creating a reusable prompt template in watsonx.ai for a marketing application. The application will pass different product names and target audiences dynamically at runtime. What is the correct syntax to define these variables within the Prompt Lab?
Show answer & explanation
Correct answer: B
In IBM watsonx.ai Prompt Lab, variables are defined using single curly braces, such as
{product_name}. When the prompt is saved as a template or used via API, these placeholders are recognized as input variables that can be dynamically populated. - Question 16Intermediate
Fine-Tuning · Understand the difference between hard and soft prompts
A data scientist needs to adapt a foundation model to classify legal documents into 50 specific categories defined by their organization. They have a dataset of 2,000 labeled examples. They want to avoid the high computational cost of retraining all model weights while still achieving high accuracy on this specific task. Which fine-tuning approach is most appropriate?
Show answer & explanation
Correct answer: B
Prompt Tuning (a form of Parameter-Efficient Fine-Tuning or PEFT) involves training a small set of learnable parameters (soft prompts) that are prepended to the input, while keeping the underlying foundation model frozen. This is highly efficient compared to full fine-tuning, requires less compute, and is ideal for classification tasks with a moderate labeled dataset (like 2,000 examples) where adapting the style/format is needed without changing the model's core knowledge.
- Question 17Intermediate
Fine-Tuning · Prepare the dataset for training
You are preparing a dataset for tuning a foundation model in watsonx.ai using the Tuning Studio. The goal is to train the model to generate marketing copy based on product features. What is the required file format for the training data?
Show answer & explanation
Correct answer: B
The watsonx.ai Tuning Studio requires training data to be in JSONL (JSON Lines) format. Each line in the file must be a valid JSON object containing
input(the prompt/context) andoutput(the desired completion) fields. While some variations exist, JSONL is the standard requirement for ingestion. - Question 18Advanced
Fine-Tuning · Customize LLMs with InstructLab
Which component of the InstructLab methodology is responsible for defining the specific skills or knowledge that you want to add to the Large Language Model?
Show answer & explanation
Correct answer: B
In InstructLab, the 'Taxonomy' is a hierarchical directory structure (YAML files) where users define the skills (tasks the model should perform) and knowledge (facts the model should know) they want to inject. Users submit contributions to the taxonomy, which are then used to generate synthetic data for training.
- Question 19Beginner
Retrieval-Augmented Generation (RAG) · Describe when to use a vector database
A solution architect is designing a RAG (Retrieval-Augmented Generation) system. They need to select a vector database to store embeddings of 10 million documents. The system requires millisecond-latency similarity searches. Which of the following is a primary function of the vector database in this architecture?
Show answer & explanation
Correct answer: B
The primary role of a vector database (like Milvus, Chroma, or Elasticsearch with vector plugins) in a RAG architecture is to store high-dimensional vectors and perform efficient similarity searches (often using algorithms like HNSW for Approximate Nearest Neighbor) to retrieve relevant context based on the query's embedding.
- Question 20Advanced
Retrieval-Augmented Generation (RAG) · Develop using libraries
When implementing a RAG solution using LangChain and watsonx.ai, you notice that the retrieved documents are not relevant to the user's specific query nuances. The retrieval is based on simple cosine similarity. What advanced technique should you implement to improve the relevance of the retrieved documents before passing them to the LLM?
Show answer & explanation
Correct answer: B
Standard vector search (bi-encoder) is fast but sometimes misses fine-grained relevance. A Re-ranking step (using a Cross-Encoder) takes the top N results from the initial retrieval and scores them more accurately against the query. This significantly improves the quality of the context provided to the LLM.
Ready for the real thing?
The full C1000-185 simulator has every exam-style question, timed mode, and instant scoring.