Knowledge Catalog overview

Knowledge Catalog is a Gemini-powered data catalog that provides universal business context and governance for your entire data estate. By automatically extracting semantics from structured and unstructured data, it builds a dynamic context graph that grounds AI agents in enterprise truth and reduces hallucinations. Data teams and AI developers use Knowledge Catalog to discover data, enforce policies, and retrieve rich context for both analytics and autonomous applications. For a detailed walkthrough of Knowledge Catalog, see the embedded video.

Dataplex Universal Catalog is now Knowledge Catalog

To better reflect the vision of unifying data governance with generative AI capabilities, Dataplex Universal Catalog is now Knowledge Catalog. This evolution of the product name represents a shift from a conventional, passive metadata registry to an active, AI-powered context graph.

Why did Dataplex become Knowledge Catalog

As organizations accelerate their generative AI adoption, AI agents need deep business context to provide accurate, grounded responses. Knowledge Catalog bridges the gap between enterprise data governance and AI agent workflows.

What is the difference between Dataplex and Knowledge Catalog

Knowledge Catalog updates reflect new AI-centric capabilities. Unlike conventional passive catalogs, Knowledge Catalog automatically curates metadata, business logic, and data relationships into a unified context graph. This graph provides the reliable enterprise truth that AI agents need to run complex tasks accurately. It leverages features like automatic context curation, verified example queries, and local and remote Model Context Protocol (MCP) integrations.

What is not changing

Your existing Dataplex deployments, APIs, and configurations remain operational. Core features like data discovery, lineage, data quality, and business glossaries are unchanged and supported. Your existing metadata, aspects, and configurations transition to the new Knowledge Catalog experience without any manual migration, data movement, or downtime.

APIs and client libraries

The rebranding to Knowledge Catalog doesn't change existing API endpoints, gcloud dataplex commands, or client libraries. You can continue to use the Knowledge Catalog APIs and client libraries to interact with Knowledge Catalog:

How Knowledge Catalog works

Knowledge Catalog unifies governance and context through three core pillars:

  • Governance foundation. Knowledge Catalog automatically collects technical metadata from Google Cloud services like BigQuery, AlloyDB for PostgreSQL, and Spanner, alongside third-party systems. It establishes a trusted data foundation through a centralized business glossary, data quality checks, anomaly detection, and policy-based governance.

  • Context curation. Using Gemini, the service infers business intent by analyzing schemas, query logs, and semantic models across your data. It generates natural language descriptions, discovers relationships, and proposes verified SQL patterns in the form of example queries that capture complex business logic.

  • Context retrieval. AI agents and applications can instantly discover assets and retrieve enriched context through semantic search and tools supporting the Model Context Protocol (MCP). This lets agents access organizational truth for reliable decision-making.

The following diagram illustrates the architecture of Knowledge Catalog and how it unifies data governance with generative AI workflows:

Architecture of Knowledge Catalog showing the curation of metadata, business logic, and data relationships into a unified context graph for AI agents. Architecture of Knowledge Catalog showing the curation of metadata, business logic, and data relationships into a unified context graph for AI agents.
Figure 1. Architecture of Knowledge Catalog (click to enlarge)

Common use cases

Knowledge Catalog helps data engineers, data scientists, and AI developers solve challenges across data management and AI development:

  • Enrich data for AI. Use data insights for unstructured data to automatically extract metadata and entities from unstructured files such as PDFs in Cloud Storage. This makes dark data and organizational knowledge accessible to AI models.

  • Reduce AI hallucinations. Provide AI agents with pre-verified example queries and semantic guardrails, letting them execute complex data retrievals with more deterministic accuracy.

  • Accelerate data discovery. Use semantic search and a centralized context graph to locate relevant data assets across disparate sources for analytics and data science workflows.

  • Automate data product creation. Infer relationships across your data estate to package assets into self-contained data products with built-in service-level agreements (SLAs) and governance constraints.

Sample workflows in Knowledge Catalog

To see how you can build your context graph and manage your data estate, consider how an online retail company might use the following Knowledge Catalog features:

  • Discover and catalog data. The retailer automatically ingests transaction data and collects metadata from Google Cloud services like BigQuery, Pub/Sub, and Cloud Storage. The service also imports metadata from custom inventory databases to build a unified view of the entire retail data estate. For more information, see Discover data.

  • Search for data assets. A data scientist finds the exact customer data assets they need using the Knowledge Catalog search engine with faceted filtering, natural language semantic search, and logical operators. For more information, see Search for data assets.