Strategic landscape of market technologies: Data Dictionary, Catalog, Lineage, Data Quality, Observability, RAG, MCP & Agentic AI
2026-03-01 · Pejman Gohari — CDO & CIO Advisory
Study overview. 7 building blocks analyzed, 30+ solutions, 12 analyst sources.
This study covers the full data governance ecosystem, expanded to include emerging building blocks driven by agentic AI. 30+ solutions were analyzed based on Gartner reports (MQ Metadata Management 2025, MQ D&A Governance 2025 and 2026), IDC MarketScape (AI Governance 2025-2026), G2 Winter 2026, Precedence Research, and public vendor data (pricing, features, roadmaps).
| Building block | Maturity | Leading solutions | Key finding |
|---|---|---|---|
| Dictionary | 6/10 | Collibra, Atlan, DataGalaxy, Databricks Unity | Average completeness around 30% across organizations. Evolving toward an active semantic layer (active metadata, +70% adoption by 2027 per Gartner). |
| Catalog | 8.5/10 | Atlan, Collibra, Alation, DataGalaxy, Databricks Unity, Informatica | Most mature building block. The differentiating criterion in 2026 is the ability to catalog AI assets (models, agents, MCP servers). Cloud Provider Catalogs (Purview, Google) cover a limited single-cloud scope. |
| Lineage | 7/10 | IBM MANTA, Atlan, Collibra, OpenLineage/Marquez, Datafold | Most underestimated and hardest to maintain building block. OpenLineage (LFAI) is emerging as the vendor-neutral standard. IBM MANTA remains the reference for complex code parsing. |
| Data Quality | 7.5/10 | Monte Carlo, Soda, Great Expectations, Bigeye, Elementary | $346M market in 2024 (+20.8% YoY, Gartner). 53% of data leaders have implemented observability (Gartner 2025). Convergence of data quality and data observability. Metaplane acquired by Datadog (Apr. 2025). |
| Agentic AI | 5/10 | Databricks Unity/Agent Bricks, Kore.ai, Salesforce Agentforce, LangChain/LangSmith | 62% of companies are experimenting, 2/3 have not deployed at scale (McKinsey). Companies practicing AI governance put 12x more projects into production (Databricks 2026). 57% of data is not AI-ready (Deloitte 2026). |
| MCP | 4.5/10 | AAIF (Linux Foundation), MCP Manager, Kong, API gateways | 10,000+ public servers. Standard transferred to AAIF (Dec. 2025, 146 members). Identified security risks: prompt injection, supply chain, tool poisoning. 2026 roadmap: auth, events, extensions. |
| RAG | 6.5/10 | Elastic, Contextual AI (RAG 2.0), Vectara, LangChain, Weaviate | $1.85B market in 2025, projected to $67.4B by 2034 (CAGR 49%, Precedence Research). 73% of implementations in large organizations. |
Rating /10 per axis, weighted: AI 25%, Governance 20%, Lineage 15%, Observability 15%, Sovereignty 15%, UX 10%.
| # | Solution | Score | Strengths | Weaknesses |
|---|---|---|---|---|
| 1 | Databricks Unity | 72.0 | AI 9/10, Lineage 9/10, Obs 8/10 | Sovereignty 4/10 |
| 2 | DataGalaxy | 69.0 | Sov 10/10, Gov 9/10, UX 8/10 | AI 4/10, Obs 5/10 |
| 3 | Atlan | 63.5 | UX 9/10 | Sov 4/10, Obs 6/10 |
| 4 | OpenMetadata | 59.0 | Balanced, Sov 7/10 (OSS) | Gov 5/10, AI 5/10, UX 5/10 |
| 5 | Collibra | 57.5 | Gov 9/10 (700+ clients) | AI 3/10, Obs 5/10, UX 5/10 |
| 6 | Informatica (Salesforce) | 56.5 | Lineage 8/10 (200+ connectors) | AI 4/10, UX 4/10, Sov 5/10 |
Detailed justifications and methodology in the Benchmark & Recommendations tab.
| Archetype | Solutions | Analyst positioning | Deployment |
|---|---|---|---|
| Enterprise | Collibra (700+ clients), Informatica/Salesforce ($8B, Nov. 2025), IBM | Leaders Gartner MQ Metadata + D&A Gov. | 3–9 months |
| Modern / Unified | DataGalaxy (200+ clients, +200% YoY, YOOI acquisition Feb. 2025) | Niche Player MQ D&A Gov. MQ Metadata G2 4.8/5 | Weeks–Months |
| Lightweight | Atlan (Leader 2 MQ 2026). Secoda → Atlassian, Select Star → Snowflake, CastorDoc → Coalesce | Atlan: Leader Metadata + D&A Gov. 2026 | Weeks |
| Lakehouse / Data Platform | Databricks Unity Catalog (Business Semantics, auto column lineage, Iceberg interop), Snowflake Horizon Catalog (+ Select Star, + Observe) | Leader IDC MarketScape 2025-2026 | Native Databricks |
| Cloud Provider | Microsoft Purview, Google Data Catalog | Cloud-native, included in subscription | Native |
| Open Source | OpenMetadata, DataHub (Acryl Data), Amundsen | LFAI / Linux Foundation | Weeks (engineering) |
| Profile | Catalog | Lineage | Quality / Observability |
|---|---|---|---|
| Large regulated enterprise | Collibra / Informatica | IBM MANTA + OpenLineage | Monte Carlo |
| Mid-market / Scale-up | Atlan / DataGalaxy | Atlan + OpenLineage | Soda + Elementary |
| Lakehouse-native | Databricks Unity | OpenLineage + Datafold | Great Expectations |
| EU Sovereignty | DataGalaxy | DataGalaxy + OpenLineage | Soda (EU-hosted) |
| SMB / Limited budget | Secoda (via Atlassian) / OpenMetadata | OpenLineage + Marquez | Soda Core + dbt tests |
| Uneven maturity | The catalog sits at 8.5/10 market maturity, MCP at 4.5/10. This gap creates risk: organizations are deploying agents on incomplete governance foundations. |
| Multi-purpose data | The same data now serves 5 layers (reporting, analytics, ML, RAG, agentic). Each layer adds quality, freshness, and semantic requirements. Without a catalog as a single source of truth, each layer silently diverges. |
| Commoditization | Standalone data discovery tools are losing their independence. In 12 months, Secoda (→ Atlassian), Select Star (→ Snowflake), CastorDoc (→ Coalesce), and Data.world (→ ServiceNow) were absorbed by horizontal platforms. The risk is no longer theoretical: with LLMs natively scanning schemas via MCP or Iceberg REST API, the value-add of the discovery "wrapper" migrates to the agent interface. Atlan remains the only independent lightweight player of note. |
| Accelerated consolidation | In 12 months, horizontal platforms (Snowflake, Atlassian, ServiceNow, Salesforce, Coalesce) absorbed 6 catalog and observability players. The "context layer" — dictionary + catalog + tribal knowledge package for agents — has become a strategic asset that every platform wants to integrate natively rather than leave to a third party. |
| Sovereignty | DataGalaxy (on-premise, self-hosted AI, multilingual) is the only European player in 2 Gartner MQs. Performance parity between sovereign models (Mistral) and frontier models (GPT-5, Claude Opus) remains an open trade-off. |
| MCP Security | The MCP protocol opens new attack surfaces: prompt injection, supply chain, tool poisoning, data exfiltration. AAIF is structuring governance, but authorization controls, DLP, and private registries remain to be implemented by each organization. |
The data governance ecosystem is undergoing a profound transformation driven by agentic AI, the MCP protocol, and tool convergence.
| Building block | Key function | Market maturity | Agentic AI impact |
|---|---|---|---|
| Data Dictionary | Semantic contract: defining business terms | Critical; an agent ingesting poorly defined data reasons on a false semantic foundation | |
| Data Catalog | Asset inventory: tables, APIs, models, dashboards | Must integrate MCP servers, RAG pipelines, model versions | |
| Lineage | Transformation traceability from source to KPI | Essential for explainability of agentic decisions | |
| Data Quality | Measuring and ensuring data reliability | AI exponentially amplifies quality errors | |
| Data Observability | Monitoring, anomaly detection, proactive alerting | 53% of data leaders have already implemented; 43% plan to within 18 months | |
| RAG (Retrieval-Augmented Generation) | Contextualizing LLMs with enterprise data | Market projected at $67B by 2034 (CAGR 49%) | |
| MCP (Model Context Protocol) | Standard for AI agent ↔ tools/data interconnection | Governed by AAIF (Anthropic, OpenAI, Block) under Linux Foundation |
The same data point, "consolidated net revenue," now flows through five simultaneous usage layers. Each layer inherits the requirements of the previous one and adds new ones. Governance must cover the entire stack — otherwise each layer silently diverges from the others.
| Usage layer | Required quality | Freshness | Semantic requirement | Consequence of error |
|---|---|---|---|---|
| Regulatory reporting | Certified, audited | Closing (D+n) | Formal (IFRS, Basel, Solvency) | Fine, sanction, license revocation |
| Operational analytics | Reliable, consistent | Daily | Contextualized for business | Poor human decision, visible in a dashboard |
| ML / Scoring | Statistically stable | Batch or streaming | Encoded, normalized, bias-free | Model drift, bias, erroneous predictions |
| RAG / GenAI | Complete, up to date | Near real-time | Indexable, chunkable, disambiguated | Contextual hallucination — false but plausible answer |
| Agentic AI | All of the above | Real-time | LLM-interpretable, unambiguous | Autonomous false decision, undetectable because syntactically correct |
The last row is the tipping point: the agent inherits all requirements from previous layers and adds a new one — machine interpretability. An incomplete dictionary bothered no one when only SQL analysts used it. When an autonomous agent consumes it, errors propagate at inference speed.
Le contenu complet de cette étude est réservé. Il est servi par l'API sur présentation d'un code — il n'est pas inclus dans les pages publiques.