AI-901 - Implement AI solutions by using Microsoft Foundry - Section 2.4

Implement AI solutions for information extraction by using Azure Content Understanding in Foundry Tools across documents, images, audio, and video.

Extracting information from documents and forms, from images, and from audio and video by using Azure Content Understanding in Foundry Tools, then building a lightweight application with information-extraction capabilities. Azure Content Understanding is the named information-extraction capability in Foundry Tools; attributing a structured-extraction task to a general vision model instead of Content Understanding is the classic single-best distractor.

Azure Content UnderstandingFoundry Toolsextract information from documents and formsextract information from imagesextract information from audio and videoinformation-extraction application

Practice question for this objective

Free sampleImplement AI solutions by using Microsoft Foundrymedium

A logistics firm receives thousands of scanned bills of lading, plus photographs of damaged crates and recorded phone calls with drivers. The firm wants a single Foundry capability that can pull structured fields and content from all of these formats: documents, images and audio. Which capability is purpose-built for this cross-format information extraction?

  • AAn image-generation model from the Foundry catalogue
  • BAzure Speech in Foundry Tools
  • CAn embedding model deployed through the Foundry SDK
  • DAzure Content Understanding in Foundry Tools Correct
Recognise Azure Content Understanding as the single Foundry Tools capability for information extraction across documents, images, audio and video. Azure Content Understanding is designed to ingest multiple modalities and return consistent structured information, which is why one capability covers documents, images and audio together rather than needing a separate service per format.

Why A is wrong: Tempting because photographs are mentioned, but an image-generation model creates new pictures from prompts rather than extracting fields from existing inputs.

Why B is wrong: Tempting because recorded calls are present, but Azure Speech handles only audio transcription and cannot extract fields from scanned documents or images.

Why C is wrong: Tempting as a way to represent mixed content, but embeddings turn input into vectors for search or similarity, not into named extracted fields.

Why D is correct: Correct. Azure Content Understanding is the named Foundry Tools capability for extracting structured information from documents, images, audio and video with one service.

See more AI-901 practice questions, answers explained.

More in this domain

Back to all Implement AI solutions by using Microsoft Foundry objectives, or the AI-901 cert hub.

Examworthy is not affiliated with or endorsed by Microsoft. Original, blueprint-aligned practice material only.