DP-600 - Prepare Data (46% of the exam) - Section 2.2

Discover data using the OneLake catalog and Real-Time hub and choose between Fabric data stores for a workload.

Use the OneLake catalog to discover and govern existing data assets, and use the Real-Time hub to find streaming data sources. Choose between a Lakehouse, Warehouse, and Eventhouse based on workload type, query pattern, and whether data is structured, semi-structured, or streaming.

OneLake catalogReal-Time hubLakehouse vs WarehouseEventhousechoosing a data store

Practice question for this objective

Free samplePrepare Datamedium

An analytics engineer is asked to connect a report to live event data and first needs to discover which streaming sources, such as existing event streams and KQL databases, are already available across the tenant so they can subscribe to one rather than configure a new feed. They want the dedicated in-product surface for browsing streaming and real-time data sources. Which feature should they open?

  • AThe "OneLake" catalog, which is the dedicated surface for browsing existing event streams and KQL databases so the engineer can subscribe to a live streaming source.
  • BThe "Warehouse" explorer, which lists existing event streams and KQL databases tenant-wide so the engineer can subscribe to a live streaming source for the report.
  • CA deployment pipeline, which lists existing event streams and KQL databases across stages so the engineer can subscribe to a live streaming source for the report.
  • DThe Real-Time hub, which lists existing event streams and KQL databases across the tenant so the engineer can discover and subscribe to a live streaming source. Correct
Use the Real-Time hub to discover existing streaming sources such as event streams and KQL databases before subscribing to live event data. The Real-Time hub is the central in-product surface for streaming and real-time data, listing event streams and KQL databases across the tenant so an engineer can find and subscribe to an existing live source instead of building a redundant feed.

Why A is wrong: The "OneLake" catalog discovers data items broadly across the tenant, but the dedicated surface for browsing streaming and real-time sources specifically is the Real-Time hub, not the catalog.

Why B is wrong: The "Warehouse" explorer browses relational tables inside one warehouse and does not catalog streaming sources, so it cannot surface event streams or KQL databases to subscribe to.

Why C is wrong: A deployment pipeline promotes content between stages of a workspace and does not catalog streaming sources, so it is not where an engineer discovers live event feeds.

Why D is correct: The Real-Time hub is the dedicated surface that centralises streaming and real-time data sources tenant-wide, letting the engineer discover existing event streams and KQL databases and subscribe to one.

See more DP-600 practice questions, answers explained.

Exam traps in Prepare Data

Answers that look right on this material and are not. Each one is a distractor from a different question in the DP-600 bank for this domain.

  • A "Warehouse", because it is the only store that exposes Delta tables to Spark notebooks and lets engineers land raw JSON and Parquet files into managed folders.

    Why it is wrong: A "Warehouse" is a T-SQL store written through SQL, not Spark, and it does not surface a file area for landing raw JSON and Parquet, so it does not fit a file-and-Spark engineering workload.

  • The "OneLake" catalog, because it is the dedicated surface that lists existing streaming sources tenant-wide so an engineer can connect to a running event stream instead of building one.

    Why it is wrong: The "OneLake" catalog discovers stored data items broadly, but the surface dedicated to browsing live streaming sources specifically is the Real-Time hub, not the catalogue.

  • Ingest the "Lakehouse" tables into the KQL database with a scheduled data pipeline so the reference data is physically resident before KQL queries run.

    Why it is wrong: Ingesting into the KQL database copies the rows and needs a scheduled job to stay current, which is exactly the duplication and overhead the requirement rules out.

Examworthy is not affiliated with or endorsed by Microsoft. Original, blueprint-aligned practice material only.