CertSafari

    Free Microsoft Certified: Azure AI Fundamentals (AI-901) Sample Questions

    35 free sample questions from our bank of 356+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Identify AI concepts and capabilities

    Subdomain 1.3: Identify AI workloads

    1.A retailer feeds thousands of product reviews into an AI service and receives a score for each review indicating whether the reviewer's opinion is positive, negative, neutral, or mixed. Which text analysis technique is being used?

    1. A.Key phrase extraction
    2. B.Sentiment analysis
    3. C.Entity recognition
    4. D.Text summarization
    Show answer & explanation

    Correct answer: B — Sentiment analysis

    • A. Key phrase extraction returns the main topics or concepts mentioned in text, not a judgment about whether the tone of the text is positive or negative.
    • B. Sentiment analysis is correct because it evaluates text and assigns a positive, negative, neutral, or mixed opinion score, which is exactly the output the retailer is receiving for each review.
    • C. Entity recognition identifies and categorizes items such as people, places, or products mentioned in text, rather than scoring the overall opinion expressed.
    • D. Text summarization condenses a document into a shorter version of its content, which does not produce a positive-or-negative opinion score.

    Subdomain 1.3: Identify AI workloads

    2.A travel app sorts a large batch of uploaded vacation photos into folders labeled "beach," "mountain," and "city," assigning one overall label to each photo. Which computer vision capability is being used?

    1. A.Object detection
    2. B.Optical character recognition (OCR)
    3. C.Sentiment analysis
    4. D.Image classification
    Show answer & explanation

    Correct answer: D — Image classification

    • A. Object detection identifies and locates multiple individual items within a single image, rather than assigning one overall category label to the whole photo.
    • B. Optical character recognition extracts printed or handwritten text from an image, which does not apply to sorting photos by overall scene type.
    • C. Sentiment analysis evaluates the tone of written text, which has no way to categorize the content of a photograph.
    • D. Image classification is correct because it assigns an overall category label, such as beach or mountain, to an entire image based on its dominant content.

    Subdomain 1.3: Identify AI workloads

    3.A team records a one-hour internal meeting and wants the recording automatically divided into timestamped chapters, each with a short recap of what was discussed during that segment. Which capability is designed for this scenario?

    1. A.Key phrase extraction
    2. B.Language identification
    3. C.Text-to-speech
    4. D.Conversation summarization
    Show answer & explanation

    Correct answer: D — Conversation summarization

    • A. Key phrase extraction returns a flat list of important topics from text, which does not divide a recording into timestamped, recapped chapters.
    • B. Language identification determines which language is present in audio, which has no role in segmenting a meeting into chapters with recaps.
    • C. Text-to-speech converts written text into spoken audio, which is the opposite of summarizing an existing meeting recording.
    • D. Conversation summarization is correct because it recaps and segments long meetings into timestamped chapters, matching exactly what the team wants from the recording.

    Subdomain 1.3: Identify AI workloads

    4.To convert a customer's recorded voicemail into written text so it can be read in a support ticket, an application would use Azure Speech's _____ feature.

    1. A.text-to-speech
    2. B.speech-to-text
    3. C.pronunciation assessment
    Show answer & explanation

    Correct answer: B — speech-to-text

    • A. Text-to-speech converts written text into spoken audio, which is the reverse of converting a recorded voicemail into text.
    • B. This is correct because speech-to-text converts recorded or streaming audio into written text, which is exactly what turns a voicemail into ticket text.
    • C. Pronunciation assessment scores how accurately a phrase was spoken; it does not produce a transcript of an entire voicemail.

    Subdomain 1.2: Identify AI model components and configurations

    5.A support team wants to explain to stakeholders why their new chatbot can produce a wrong but confident-sounding answer. Which explanation best describes how the underlying generative model produces its response?

    1. A.It repeatedly predicts the most probable next token from patterns learned in training, with no built-in fact-checking step.
    2. B.It retrieves a pre-written answer from a curated internal database, so any mistake must come from a single incorrect database entry.
    3. C.It follows a fixed rule-based decision tree authored by developers, so an error means one specific rule was written incorrectly.
    4. D.It runs a live web search on every incoming request and copies the top result, so an error means the search index itself was stale.
    Show answer & explanation

    Correct answer: A — It repeatedly predicts the most probable next token from patterns learned in training, with no built-in fact-checking step.

    • A. Generative language models generate text by predicting the next most likely token based on patterns learned from massive training data, one token at a time. This process does not include a verification step against ground truth, which is why fluent, confident-sounding output can still be factually wrong.
    • B. This describes a lookup or retrieval system, not how a generative model works. Generative models synthesize new text token by token rather than returning a stored, pre-written answer.
    • C. This describes a rule-based expert system, not a generative model. Generative models rely on learned statistical patterns from training data rather than explicit if-then rules written by a developer.
    • D. This describes retrieval-augmented behavior with a live search backend, not the base mechanism of a generative model. A standalone generative model produces text from learned patterns, not a real-time web lookup, unless it has been explicitly connected to a search tool.

    Subdomain 1.2: Identify AI model components and configurations

    6.What is a token in the context of a large language model?

    1. A.A chunk of text, such as a whole word or part of a word, that the model treats as its smallest processing unit.
    2. B.A single training example consisting of one complete labeled sentence used once during the model's original training run.
    3. C.A unique numeric identifier assigned to each user session that keeps track of conversation history across requests.
    4. D.A security credential issued by Azure Foundry that authorizes an application to call a deployed model endpoint.
    Show answer & explanation

    Correct answer: A — A chunk of text, such as a whole word or part of a word, that the model treats as its smallest processing unit.

    • A. Language models break input and output text into tokens, which can be whole words, parts of words, or punctuation, and process one token at a time when predicting what comes next.
    • B. This describes a labeled training example, not the unit of text the model reads and generates during inference.
    • C. This describes a session identifier used by an application, which is unrelated to how a model represents text internally.
    • D. This describes an authentication credential for calling an API, not the unit of text a language model processes.

    Subdomain 1.2: Identify AI model components and configurations

    7.A learner asks why a generative model can write a plausible poem about a topic it was never explicitly trained on with that exact wording. What is the best conceptual explanation?

    1. A.It learned general language patterns from a massive training corpus, which lets it generalize to new word combinations it never saw verbatim.
    2. B.The model stores every training sentence in a searchable index and reassembles the closest matching fragments to answer each new request.
    3. C.The model contacts a live internet search engine during each generation step and paraphrases the closest matching poem it can find.
    4. D.The model was manually pre-programmed with a fixed response template that was written in advance for every possible topic a user might request.
    Show answer & explanation

    Correct answer: A — It learned general language patterns from a massive training corpus, which lets it generalize to new word combinations it never saw verbatim.

    • A. Training on a massive text corpus teaches a model statistical patterns of language, letting it generate new, coherent combinations of words for topics it never saw phrased that way before.
    • B. Generative models do not store and retrieve verbatim training sentences from an index; they generate new token sequences from learned patterns instead.
    • C. A base generative model does not perform a live web search during generation unless it has been explicitly connected to a search tool.
    • D. Generative models do not rely on manually written templates for each topic; their flexibility comes from patterns learned during training.

    Subdomain 1.2: Identify AI model components and configurations

    8.A machine learning lead is drafting selection criteria before the team commits to a model from the catalog for a new customer-facing assistant. Which of the following are reasonable criteria to include? (Select all that apply)(Select 3)

    1. A.Whether the model's context window is large enough to hold the assistant's typical conversation length.
    2. B.Whether the model's licensing terms and support model fit the team's compliance and support needs.
    3. C.Whether the model's benchmark results on relevant tasks meet the assistant's accuracy requirements.
    4. D.Whether the model card's illustration uses the same color scheme as the company's brand guide.
    5. E.Whether the model was the very first entry added to the catalog when it first launched.
    Show answer & explanation

    Correct answers: A, B, C — Whether the model's context window is large enough to hold the assistant's typical conversation length.; Whether the model's licensing terms and support model fit the team's compliance and support needs.; Whether the model's benchmark results on relevant tasks meet the assistant's accuracy requirements.

    • A. Matching context window size to expected conversation length prevents truncation of important earlier context, making it a reasonable selection criterion.
    • B. Licensing and support terms differ between Azure-sold and partner/community models and affect long-term reliability, making this a legitimate criterion.
    • C. Benchmark performance on tasks similar to the assistant's use case is a direct, practical way to judge whether a model will meet accuracy needs.
    • D. A model card's visual color scheme has no bearing on the model's technical fit for a customer-facing assistant.
    • E. How early a model was added to the catalog says nothing about its current suitability, benchmarks, or licensing for a new project.

    Subdomain 1.2: Identify AI model components and configurations

    9.A team choosing between two similarly accurate models for a latency-sensitive feature should give extra weight to each model's ___ before making a final decision.

    1. A.cost per token and typical response latency
    2. B.the logo design shown on its model card
    3. C.its alphabetical position in the model catalog
    Show answer & explanation

    Correct answer: A — cost per token and typical response latency

    • A. When accuracy is comparable, cost per token and typical response latency become the meaningful differentiators for a latency-sensitive feature.
    • B. A model card's logo design has no bearing on the model's suitability for a latency-sensitive feature.
    • C. Where a model's name falls alphabetically in the catalog has no relationship to its cost, latency, or fit for the task.

    Subdomain 1.1: Describe principles of responsible AI

    10.A hospital deploys an AI model that suggests a likely diagnosis from patient symptoms and test results. To follow responsible AI guidance, the system is designed so that whenever its confidence score falls below a set threshold, it withholds a suggestion and instead flags the case for a physician to review manually. Which principle does this design choice primarily support?

    1. A.Reliability and safety, because the system fails safely rather than presenting a low-confidence diagnosis as certain
    2. B.Privacy and security, because withholding the suggestion keeps the patient's diagnostic data from unauthorized parties
    3. C.Inclusiveness, because physicians from a wide range of specialties can all understand the flagged cases equally well
    4. D.Fairness, because every patient's case is flagged for review using exactly the same confidence threshold value
    Show answer & explanation

    Correct answer: A — Reliability and safety, because the system fails safely rather than presenting a low-confidence diagnosis as certain

    • A. Withholding a low-confidence output and deferring to a human reviewer is a classic reliability and safety pattern: the system avoids presenting an unreliable result as though it were dependable.
    • B. The design choice is about output confidence and safe failure, not about who can access patient data, so it does not primarily address privacy and security.
    • C. Physician understanding across specialties is not what the threshold mechanism controls; the mechanism controls when the model defers rather than how clearly its output is explained.
    • D. Applying one fixed threshold to every case is a design detail, not evidence of equitable treatment across demographic groups, which is what the fairness principle addresses.

    Subdomain 1.1: Describe principles of responsible AI

    11.A generative AI writing assistant is trained primarily on documents written in formal, corporate English. Employees who are non-native English speakers or who write in a more conversational style report that the assistant frequently rewrites their text in ways that erase their natural voice and suggests fewer improvements are needed for text that already matches the formal corporate style. Which principle does this pattern most directly raise a concern about?

    1. A.Fairness, because the assistant's suggestions systematically favor writing styles associated with one group of employees over others
    2. B.Privacy and security, because the assistant stores drafts of employee writing in a shared, unencrypted database accessible to all staff
    3. C.Reliability and safety, because the assistant occasionally produces a sentence with a minor grammatical error
    4. D.Transparency, because employees were never told which underlying language model powers the writing assistant
    Show answer & explanation

    Correct answer: A — Fairness, because the assistant's suggestions systematically favor writing styles associated with one group of employees over others

    • A. The assistant treating writing associated with one style or group more favorably than another is a disparate-impact pattern, which is precisely the concern the fairness principle addresses.
    • B. Where drafts are stored is a data-protection question, but the scenario's concern is about unequal treatment of writing styles, not about who can access stored drafts.
    • C. Occasional grammar mistakes are a general quality issue, but the scenario specifically describes systematically unequal treatment by writing style, which is a fairness concern rather than a robustness one.
    • D. Not knowing which model powers the assistant is a disclosure gap, but the scenario's central concern is unequal suggestion quality by group, which is a fairness issue.

    Subdomain 1.1: Describe principles of responsible AI

    12.Before launching an AI-powered resume-screening tool, which combination of steps best addresses responsible AI risk across multiple principles at once? (Select all that apply.)(Select 3)

    1. A.Test screening outcomes across gender, age, and disability status to check for disparate impact patterns
    2. B.Publish a short notice telling candidates their resume will first be screened by an automated AI system
    3. C.Require a recruiter to review and be able to override any AI rejection before a candidate is removed
    4. D.Reduce the maximum resume file size that the online upload form is willing to accept from candidates
    5. E.Change the color scheme of the careers website so it matches the company's new branding guidelines
    Show answer & explanation

    Correct answers: A, B, C — Test screening outcomes across gender, age, and disability status to check for disparate impact patterns; Publish a short notice telling candidates their resume will first be screened by an automated AI system; Require a recruiter to review and be able to override any AI rejection before a candidate is removed

    • A. Testing outcomes across gender, age, and disability status checks for disparate impact on protected groups, which is a fairness evaluation appropriate before launch.
    • B. Notifying candidates that an AI system performs an initial screening is a transparency disclosure that helps candidates understand the process.
    • C. Giving a recruiter override authority over AI rejections establishes human oversight of a high-impact decision, which is an accountability safeguard.
    • D. Limiting upload file size is a technical form constraint unrelated to fairness, transparency, or accountability risks in the screening decision.
    • E. A website color scheme change is a branding decision unrelated to any responsible AI principle relevant to the screening tool.

    Subdomain 1.1: Describe principles of responsible AI

    13.Once an AI system has automated content filters in production to block harmful outputs, human oversight of the system's decisions is no longer necessary for that system to satisfy the accountability principle.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. This is a true statement of the flawed claim being tested, so True would be the wrong choice: automated filters reduce some risks but do not replace the need for human oversight and responsibility, which accountability still requires.
    • B. This statement is correctly False: content filters are a mitigation layer, but accountability still requires that people remain responsible for the system's behavior and retain the ability to intervene, even with automated filters in place.

    Subdomain 1.1: Describe principles of responsible AI

    14.Running an AI-powered chatbot through stress tests that simulate network outages and unexpected user inputs before launch supports the ___ principle.

    1. A.privacy and security
    2. B.accountability
    3. C.reliability and safety
    Show answer & explanation

    Correct answer: C — reliability and safety

    • A. Privacy and security concerns protecting data and access, not testing the system's behavior under outages and unexpected inputs, so it does not fit this description.
    • B. Accountability concerns governance and human responsibility for the system, not stress-testing its technical robustness, so it does not fit this description.
    • C. Simulating outages and unexpected inputs before launch is edge-case and stress testing aimed at uncovering failure modes, which is the reliability and safety principle.

    Subdomain 1.3: Identify AI workloads

    15.A finance team uploads scanned vendor invoices and wants an AI service to return the vendor name, invoice date, and total due as structured key-value fields ready to load into their accounting system. Which AI workload fits this need?

    1. A.Speech translation converts vendor audio into text.
    2. B.Sentiment analysis scores emotional tone in the text.
    3. C.Content Understanding extracts fields from scans.
    4. D.Image classification assigns a category label to the page.
    Show answer & explanation

    Correct answer: C — Content Understanding extracts fields from scans.

    • A. Speech translation converts spoken audio from one language into another, which has no role in pulling structured fields out of a scanned document.
    • B. Sentiment analysis scores the tone of text, which does not produce structured fields like vendor name or invoice date.
    • C. Information extraction with Content Understanding is correct because it defines a field schema and extracts structured key-value data, such as vendor name and total due, from documents like invoices.
    • D. Image classification assigns a single overall label to a whole image, such as "invoice," without extracting the individual fields inside it.

    Domain 2: Implement AI solutions by using Microsoft Foundry

    Subdomain 2.2: Implement AI solutions for text and speech by using Foundry

    16.A developer wants a multimodal model deployed in Foundry to answer questions that a user asks by speaking into a microphone, rather than typing. Which capability of the deployed model makes this possible without a separate speech-to-text step in the application code?

    1. A.The model accepts audio as an input modality alongside or instead of text in the prompt
    2. B.The model automatically increases its context window whenever audio is detected
    3. C.The model switches to a lower temperature setting when it receives a spoken prompt
    4. D.The model requires the developer to pre-encode audio as base64 key phrases before sending it
    Show answer & explanation

    Correct answer: A — The model accepts audio as an input modality alongside or instead of text in the prompt

    • A. A multimodal model that supports audio input can consume spoken prompts directly, removing the need for the application to run a separate speech-to-text transcription step first.
    • B. Context window size is a fixed property of the model deployment and does not change automatically based on whether the input happens to include audio.
    • C. Temperature controls the randomness of generated text and is set explicitly by the caller; it is not something the model silently changes based on input type.
    • D. Key phrases are an output of text analysis, not an audio encoding format, so this describes a step that does not exist in a multimodal audio workflow.

    Subdomain 2.2: Implement AI solutions for text and speech by using Foundry

    17.A developer wants a synthesized voice to pause briefly between two sentences and pronounce an abbreviation phonetically, features that plain `speak_text_async` calls cannot express. What should the app send to the synthesizer instead?

    1. A.An SSML document passed to `speak_ssml_async`, using tags to control pauses and pronunciation
    2. B.A longer plain-text string with extra punctuation marks inserted between the sentences
    3. C.A `SpeechConfig` object with `speech_synthesis_voice_name` set to a different neural voice
    4. D.A JSON payload sent to the Language service's `analyze_sentiment` endpoint before synthesis
    Show answer & explanation

    Correct answer: A — An SSML document passed to `speak_ssml_async`, using tags to control pauses and pronunciation

    • A. SSML lets a developer mark up text with tags for pauses, phonetic pronunciation, prosody, and more, and passing that markup to `speak_ssml_async` gives fine control that plain text calls cannot provide.
    • B. Extra punctuation may slightly change timing but cannot reliably force a pause or specify phonetic pronunciation the way explicit SSML tags do.
    • C. Changing the voice name selects a different synthesized voice but does not add pause control or phonetic pronunciation for a specific abbreviation.
    • D. Sentiment analysis scores text for tone and has no role in controlling how the Speech service pronounces or paces synthesized audio.

    Subdomain 2.2: Implement AI solutions for text and speech by using Foundry

    18.A voice-name string such as `en-US-AvaNeural` identifies both a locale and a specific synthesized voice, and setting it changes only how text-to-speech sounds, not how speech-to-text recognizes audio.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. Voice names combine a locale prefix with a specific voice identifier and are consumed by `speech_synthesis_voice_name`, which affects only text-to-speech output and has no effect on the separate speech-to-text recognition path, matching this statement.
    • B. This option would claim the voice name setting also affects recognition, which contradicts the documented separation between the synthesis voice setting and the recognition language setting.

    Subdomain 2.2: Implement AI solutions for text and speech by using Foundry

    19.`AudioConfig` can be constructed with either `use_default_microphone=True` for live capture or `filename="..."` to read from a pre-recorded audio file.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. Both constructor options are documented ways to configure a `SpeechRecognizer`'s audio source, giving the developer a choice between live microphone capture and reading from a file, which matches this statement.
    • B. This option would claim only one of the two input options is actually supported, which contradicts the documented flexibility of `AudioConfig` to accept either a microphone flag or a filename.

    Subdomain 2.2: Implement AI solutions for text and speech by using Foundry

    20.A developer builds a moderation tool that scans user-submitted comments for phone numbers and government ID numbers before the comments are published, and separately checks whether each comment's tone is hostile. Which two Text Analytics methods together cover this requirement? (Select 2)(Select 2)

    1. A.`recognize_pii_entities(documents=comments)`, to flag phone numbers and ID numbers
    2. B.`analyze_sentiment(documents=comments)`, to check whether the tone is hostile
    3. C.`extract_key_phrases(documents=comments)`, to flag phone numbers and ID numbers
    4. D.`detect_language(documents=comments)`, to check whether the tone is hostile
    5. E.`extractive_summarize(documents=comments)`, to flag phone numbers and ID numbers
    Show answer & explanation

    Correct answers: A, B — `recognize_pii_entities(documents=comments)`, to flag phone numbers and ID numbers; `analyze_sentiment(documents=comments)`, to check whether the tone is hostile

    • A. Correct: `recognize_pii_entities` is designed to detect sensitive personal data such as phone numbers and government ID numbers, directly matching the moderation tool's first requirement.
    • B. Correct: `analyze_sentiment` returns a tone label, including negative sentiment, which the moderation tool can use to flag hostile-sounding comments.
    • C. Incorrect: key phrase extraction returns general topical phrases, not categorized sensitive data like phone numbers or ID numbers, so it does not fulfill the PII-flagging requirement.
    • D. Incorrect: language detection identifies which language a comment is written in and produces no tone information, so it cannot fulfill the hostility-check requirement.
    • E. Incorrect: extractive summarization pulls important sentences to form a condensed excerpt; it does not tag or flag sensitive personal data like phone or ID numbers.

    Subdomain 2.1: Implement generative AI apps and agents by using Foundry

    21.A developer testing a model in the Foundry portal playground notices the assistant keeps answering in a formal tone even though the developer never asked for formality in the current message. What most likely explains this?

    1. A.A system prompt configured earlier in the session is instructing the model to maintain a formal tone
    2. B.The model automatically switches to formal tone whenever the conversation exceeds three turns
    3. C.The deployed model's temperature parameter was set below zero, which forces formal wording
    4. D.Foundry silently appends a hidden formality flag to every single user message before it reaches the model
    Show answer & explanation

    Correct answer: A — A system prompt configured earlier in the session is instructing the model to maintain a formal tone

    • A. A system prompt persists across the whole session, so an earlier instruction to stay formal continues to shape every reply the model produces.
    • B. Turn count has no built-in effect on tone; Foundry does not switch behavior automatically based on how many messages have been exchanged.
    • C. Temperature controls response randomness and cannot go below zero, and it has no mechanism that forces a formal writing style.
    • D. Foundry does not inject hidden formality flags into user messages; any consistent tone comes from prompt content the developer supplied.

    Subdomain 2.1: Implement generative AI apps and agents by using Foundry

    22.A company builds an internal HR assistant on Foundry. The system prompt says: "Only answer questions about vacation policy, benefits, and payroll dates. Decline all other topics." An employee then asks, "Can you also help me debug my Python script?" What should happen?

    1. A.The model should decline the coding request because it falls outside the scope the system prompt defines
    2. B.The model should answer the coding request because user prompts always override system prompt restrictions
    3. C.The model should switch its deployed model version automatically to one suited for coding
    4. D.The model should ignore the system prompt entirely once any off-topic question is asked
    Show answer & explanation

    Correct answer: A — The model should decline the coding request because it falls outside the scope the system prompt defines

    • A. The system prompt's scope restriction persists for the whole conversation, so an off-topic coding request should be declined as instructed.
    • B. System prompt instructions are meant to constrain behavior across the session; a user's request does not automatically override those standing rules.
    • C. A conversation cannot trigger a change of the underlying deployed model; the model version is fixed by the deployment, not by prompt content.
    • D. One off-topic question does not cause Foundry to discard the system prompt; the instruction continues to apply unless it is explicitly changed.

    Subdomain 2.1: Implement generative AI apps and agents by using Foundry

    23.In the Foundry portal, after a model finishes deploying, the developer is automatically moved into the ___ section with the new model already selected for interactive testing.

    1. A.Build
    2. B.Cost Management
    3. C.Access Control (IAM)
    4. D.Diagnostic Settings
    Show answer & explanation

    Correct answer: A — Build

    • A. The Build section is where the portal lands the developer once deployment completes, with the model ready to try in the playground.
    • B. Cost Management tracks spending on resources and is not where a freshly deployed model becomes available for testing.
    • C. Access Control (IAM) manages permissions and role assignments, not model deployment or testing.
    • D. Diagnostic Settings configures logging destinations and plays no role in presenting a newly deployed model for testing.

    Subdomain 2.1: Implement generative AI apps and agents by using Foundry

    24.A junior developer asks why they should test a model in the Foundry portal playground before writing any client code. Which two reasons best explain this practice?(Select 2)

    1. A.It confirms the model responds as expected before time is spent building a full application
    2. B.It lets the developer refine prompt wording quickly without editing and rerunning a script
    3. C.It is the only way Foundry ever allows a deployed model to be called at all
    4. D.It permanently locks the model's temperature and other parameters for future client calls
    5. E.It replaces the need to configure authentication when the client application is built later
    6. F.It automatically writes a production-ready application from the playground session
    Show answer & explanation

    Correct answers: A, B — It confirms the model responds as expected before time is spent building a full application; It lets the developer refine prompt wording quickly without editing and rerunning a script

    • A. A quick playground check is a sanity test on the model's behavior, catching problems early before code and integration effort are spent.
    • B. Editing a prompt in the playground and resending it is faster than changing a script and rerunning it for every small wording change.
    • C. The playground is one convenient way to interact with a model, but client code and the REST API can also call the deployed model directly.
    • D. Trying settings in the playground does not lock those values for every future call; client code can send its own parameters independently.
    • E. Playground testing does not set up authentication for a separate client application; that application still needs its own credential configuration.
    • F. The playground supports interactive testing only; it does not generate a deployable client application from that session.

    Subdomain 2.1: Implement generative AI apps and agents by using Foundry

    25.A developer is writing a Python chat client using the Foundry SDK. Which credential class from azure-identity is used in the quickstart samples to authenticate the AIProjectClient?

    1. A.DefaultAzureCredential
    2. B.ManagedIdentityOnlyCredential
    3. C.StaticApiKeyCredential
    4. D.LegacyBasicAuthCredential
    Show answer & explanation

    Correct answer: A — DefaultAzureCredential

    • A. DefaultAzureCredential is the credential class the Foundry SDK samples use, and it works after signing in with az login for local development.
    • B. There is no such class as ManagedIdentityOnlyCredential in azure-identity; it is not the credential used in these Foundry SDK samples.
    • C. Foundry SDK samples do not authenticate with a static API key class; they rely on Microsoft Entra ID credentials instead.
    • D. There is no LegacyBasicAuthCredential class in azure-identity, and Foundry chat client samples do not use basic authentication.

    Subdomain 2.4: Implement AI solutions for information extraction by using Foundry

    26.A retail company has field staff photograph nutrition labels on product packaging with their phones and wants the ingredient list and serving size pulled into a product database automatically. What should they configure?

    1. A.A video analyzer configured for speaker recognition, since it processes visual content the same way as still images.
    2. B.A document analyzer that only accepts native PDF or Word files, since photos cannot be processed for field extraction.
    3. C.A Content Understanding analyzer with a field schema for ingredients and serving size, applied directly to the label photos.
    4. D.A generic image resizing tool with no defined schema, since Content Understanding cannot read text inside photos.
    Show answer & explanation

    Correct answer: C — A Content Understanding analyzer with a field schema for ingredients and serving size, applied directly to the label photos.

    • A. Speaker recognition is an audio and video capability for distinguishing between people talking in a recording, which is unrelated to reading printed text on a photographed label. It does not extract field values from a still image.
    • B. Content Understanding processes images directly as one of its supported modalities, so restricting input to native PDF or Word files is unnecessary and incorrect. Photos are a valid input for field extraction.
    • C. An analyzer with a field schema for ingredients and serving size, applied directly to the label photos, extracts exactly the values the retail company needs from the image content. This is the correct configuration because it matches the input type and the required fields.
    • D. A resizing tool with no schema does not extract any field values at all, and the premise that the service cannot read text in photos is incorrect since Content Understanding processes image content directly. This does not meet the stated goal of pulling structured fields.

    Subdomain 2.4: Implement AI solutions for information extraction by using Foundry

    27.Content Understanding defines three ways to generate a field's value from content. Which set correctly names all three?

    1. A.Extract, Classify, and Generate
    2. B.Translate, Summarize, and Redact
    3. C.Detect, Segment, and Archive
    4. D.Encrypt, Compress, and Transcribe
    Show answer & explanation

    Correct answer: A — Extract, Classify, and Generate

    • A. Extract, Classify, and Generate are the three field extraction methods Content Understanding defines for producing a field's value, covering literal copying, category assignment, and free-form generation respectively. This is the correct set of method names.
    • B. Translate, Summarize, and Redact are not the names Content Understanding uses for its field extraction methods, even though summarizing content resembles part of what the Generate method can do. This set does not match the defined terminology.
    • C. Detect, Segment, and Archive are not the three defined field extraction methods; segmentation is a separate pipeline component, and detect and archive are not method names the service uses for field values. This set is incorrect.
    • D. Encrypt, Compress, and Transcribe describe unrelated technical operations rather than the field extraction methods Content Understanding defines. Transcription is part of audio and video processing, but it is not one of the three named field extraction methods.

    Subdomain 2.4: Implement AI solutions for information extraction by using Foundry

    28.A team is automating order processing with Content Understanding output feeding directly into an RPA workflow. Which practices fit that goal? (Select 3.)(Select 3)

    1. A.Rebuilding the extraction schema by hand for every single order that arrives, even when the layout repeats
    2. B.Defining a field schema so extracted order data maps directly to the fields the RPA workflow expects
    3. C.Ignoring confidence scores entirely and auto-approving every extracted order without exception
    4. D.Using confidence scores to auto-approve high-certainty orders and route uncertain ones to a person
    5. E.Requiring a person to retype every field manually before the RPA workflow can begin
    6. F.Using grounding data so a reviewer can quickly check a flagged field against the source document
    Show answer & explanation

    Correct answers: B, D, F — Defining a field schema so extracted order data maps directly to the fields the RPA workflow expects; Using confidence scores to auto-approve high-certainty orders and route uncertain ones to a person; Using grounding data so a reviewer can quickly check a flagged field against the source document

    • A. Rebuilding the schema by hand for every order, even when the layout repeats, wastes effort that a single reusable schema definition already covers. This works against efficient automation rather than supporting it.
    • B. Defining a field schema that maps directly to the fields the RPA workflow expects lets extracted order data flow straight into that workflow without extra transformation. This is one of the three correct practices because it directly connects extraction output to the automation.
    • C. Ignoring confidence scores and auto-approving every order removes the safeguard that catches unreliable extractions, increasing the risk of processing incorrect order data. This does not fit a well-designed automated workflow.
    • D. Using confidence scores to auto-approve high-certainty orders while routing uncertain ones to a person balances automation speed with accuracy safeguards. This is one of the three correct practices because it uses the reliability signal as intended.
    • E. Requiring a person to manually retype every field before the workflow begins reintroduces the manual effort that automating with Content Understanding is meant to remove. This does not fit the goal of feeding output directly into RPA.
    • F. Grounding data lets a reviewer quickly check a flagged field against its location in the source document, speeding up the manual review step when it is needed. This is one of the three correct practices because it supports efficient verification.

    Subdomain 2.4: Implement AI solutions for information extraction by using Foundry

    29.A call center can use a Content Understanding audio analyzer to both transcribe a recorded call and generate a structured sentiment field in the same processing step.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. An audio analyzer can transcribe a recording and generate structured fields, such as sentiment, together as part of one analysis, which matches how Content Understanding combines transcription with field extraction. This statement accurately describes that combined capability.
    • B. This would be incorrect because combining transcription with generated structured fields in a single audio analysis pass is exactly how the service is designed to process call recordings.

    Subdomain 2.4: Implement AI solutions for information extraction by using Foundry

    30.A media company wants a one-paragraph description of the action happening in each video scene created automatically, rather than lifted verbatim from any single frame. This uses the ____ field extraction method.

    1. A.Extract
    2. B.Classify
    3. C.Generate
    Show answer & explanation

    Correct answer: C — Generate

    • A. Extract copies a value as it already appears in the content, which does not fit writing a new paragraph describing action that is not printed anywhere in the video.
    • B. Classify assigns content to one of a fixed set of categories rather than composing a free-form written description of the action in a scene.
    • C. Generate creates a new field value from the content, which fits writing a one-paragraph description of a scene's action that is not copied from any single frame.

    Subdomain 2.3: Implement AI solutions with computer vision and image-generation capabilities by using Foundry

    31.In the Foundry portal Images playground, a user submits the prompt "a red bicycle leaning against a brick wall, photorealistic" and requests two output images. Which two settings from the playground directly shape that request? (Select 2)(Select 2)

    1. A.The system prompt that defines an agent's persistent instructions
    2. B.The maximum token budget allowed for the assistant's next reply
    3. C.The temperature used for text-completion sampling in a chat
    4. D.The output image size or aspect ratio for generated images
    5. E.The number of images to generate for each single prompt
    Show answer & explanation

    Correct answers: D, E — The output image size or aspect ratio for generated images; The number of images to generate for each single prompt

    • A. A system prompt configures a conversational agent's ongoing behavior across turns; the Images playground issues single, stateless image-generation calls with no agent persona involved.
    • B. A max-token budget limits how much text a chat model can write in a reply; it has no bearing on an image-generation request that returns pictures, not text.
    • C. Temperature governs randomness in text token sampling for chat completion models; the image generation endpoint used here does not sample text tokens in this way.
    • D. The Images playground also exposes an output size or aspect-ratio setting, which controls the dimensions of the generated bicycle image and is part of the same request.
    • E. The Images playground exposes a count setting controlling how many image variations are returned for one prompt, directly matching the request for two outputs.

    Subdomain 2.3: Implement AI solutions with computer vision and image-generation capabilities by using Foundry

    32.A user asks a deployed vision-enabled model to identify a person by name from an uploaded photo of a crowd at a public event. Which responsible-AI concern is most directly raised?

    1. A.Privacy, since identifying individuals from a photo risks processing personal data
    2. B.Reliability and safety, because the photo file size exceeds a typical upload limit
    3. C.Inclusiveness, because the request omits accessibility features for disabled users here
    4. D.Transparency, because the model was not asked to disclose its training data sources here
    Show answer & explanation

    Correct answer: A — Privacy, since identifying individuals from a photo risks processing personal data

    • A. Trying to identify named individuals from a crowd photo is exactly the kind of personal-data processing the privacy and security principle warns about, since it risks exposing or misusing people's identities without consent.
    • B. Reliability and safety concerns how consistently and safely a system performs its intended function; a file-size limit is an engineering constraint, not the responsible-AI issue this scenario is testing.
    • C. Inclusiveness is about designing systems that work well for people of varying abilities and backgrounds; it does not describe the risk of identifying individuals in a photo without consent.
    • D. Transparency is about disclosing that an AI system is involved and explaining its outputs to users; disclosing training-data sources in a single reply is not the concern this scenario raises.

    Subdomain 2.3: Implement AI solutions with computer vision and image-generation capabilities by using Foundry

    33.A team wants a vision-enabled chat model to describe an uploaded photo, and the client code must place the encoded image inside a content part whose type is set to the blank below. Complete the statement: In the chat message, the image is added as a content part with type ____.

    1. A.`vision_content`
    2. B.`image_url`
    3. C.`image_file`
    Show answer & explanation

    Correct answer: B — `image_url`

    • A. This is not the field name defined by the API; the actual schema uses `image_url` as the content-part type for image input.
    • B. This is the documented content-part type in the chat completions message schema used to carry either a hosted URL or a base64 data URI for an image.
    • C. This is not a recognized content-part type in the chat completions message schema; images are carried through the `image_url` part regardless of whether the source is local or hosted.

    Subdomain 2.3: Implement AI solutions with computer vision and image-generation capabilities by using Foundry

    34.A real-estate app lets an agent photograph an empty room and generate a staged version showing furniture in a chosen style. Which combination best fits this workflow?

    1. A.Send the photo to a speech-to-text service so the agent can narrate the room aloud
    2. B.Send the photo to a sentiment analysis client to score how inviting the room appears
    3. C.Send the photo to a key-phrase extraction client to list nouns describing the room
    4. D.Send the room photo to the image model's edits capability with a staging prompt
    Show answer & explanation

    Correct answer: D — Send the room photo to the image model's edits capability with a staging prompt

    • A. Speech-to-text converts spoken audio into text; it does not analyze or generate images and has no role in producing a staged room photo.
    • B. Sentiment analysis scores emotional tone in text and cannot process or alter an image, so it cannot produce a staged photo at all.
    • C. Key-phrase extraction pulls topical phrases from written text; it cannot process image pixels or generate a new staged photo.
    • D. Passing the existing empty-room photo into the image-generation model's edits capability, with a prompt describing the desired furniture and style, lets the model add staged furniture while preserving the room's original structure.

    Subdomain 2.3: Implement AI solutions with computer vision and image-generation capabilities by using Foundry

    35.A user asks a deployed multimodal model to interpret a photo of a medical X-ray and suggest a diagnosis for a home health app with no clinician involved. Which responsible-AI principle is most at risk if the app ships this way?

    1. A.Fairness, since the X-ray was taken with a smartphone camera instead of clinical equipment here
    2. B.Reliability and safety, since acting on an unverified reading with no oversight risks real harm
    3. C.Inclusiveness, since the app lacks multiple spoken languages for its diagnosis text
    4. D.Transparency, because the reply is returned as plain text rather than structured JSON
    Show answer & explanation

    Correct answer: B — Reliability and safety, since acting on an unverified reading with no oversight risks real harm

    • A. Camera equipment quality could affect image clarity, but the principal responsible-AI concern here is the lack of human medical oversight before acting on the output, which is a safety issue rather than a fairness one.
    • B. Using an AI-generated reading of a medical image to drive a diagnosis with no clinician review is a direct reliability and safety risk, since an incorrect or overconfident output could lead to real patient harm.
    • C. Multi-language support is about accessibility and reach, not about the core risk in this scenario, which is acting on an unverified medical judgment without qualified oversight.
    • D. Returning plain text instead of structured JSON is a formatting choice, not a transparency issue about disclosing AI involvement or explaining the output to the user.

    Want the full experience?

    These are just samples. Practice the full Microsoft Certified: Azure AI Fundamentals (AI-901) question bank in quiz mode — free, no signup, with domain practice and exam simulation.