Conduktor is the governance hub for streaming data and AI. We commissioned this report to understand where organizations are succeeding, and struggling, with turning streaming data into governed pipelines that improve analytics and decision-making, accelerate AI initiatives, and increase collaboration with external partners.
One key conclusion: while most respondents report greater use of data streaming, particularly to feed AI systems with fresh data from multiple sources, the proliferation of streaming platforms creates big problems. Those problems include losing data when consolidating platforms and challenges integrating between systems. In the race to get ahead in AI, organizations cannot ignore the need for strong governance, and risk corporate-wide chaos when they do.
93%use multi-platform data streaming architectures
43%already use data streaming for training or running AI
86%struggle with integration across streaming platforms
98%are concerned about losing data when consolidating
The report is based on independent research from Pureprofile with 200 senior IT and data executives across Europe and the US. Respondents were drawn from companies with an annual revenue of $50 million or more and over 500 employees, operating in financial services, manufacturing, retail, telecoms, transportation and logistics, and utilities.
200senior IT & data executives surveyed
500+employees (company size)
$50M+annual revenue
The core tension revealed by our research is that while streaming adoption is accelerating, so too are the complexity and fragmentation that emerge as organizations adopt more platforms.
Use multi-platform architectures93%
Use more platforms than two years ago77%
Consolidated down to one platform16%
Use fewer platforms than two years ago7%
Larger organizations in particular are consolidating platforms to reduce architectural complexity, lower costs, and cut governance overhead. But for most, the direction of travel is more platforms, not fewer, as teams scale streaming to address new use cases.
The picture so far
- Streaming is now multi-platform by default. 93% run more than one platform, and most are adding rather than consolidating.
- Fragmentation is the side effect. Each new platform brings its own governance model, integration surface, and cost profile.
- AI is the accelerant. The rest of this report looks at how teams are feeding AI with streaming data, where it breaks, and what closes the gap.
The adoption of AI is now firmly at the heart of strategic IT and business planning, and data streaming has become a vital element of it, channeling relevant information from multiple sources to decision-makers and models.
Automating workflows83%
Making real-time decisions51%
Training or running AI43%
The top three areas where respondents say they've benefited most from streaming are enhancing customer experience with faster, personalized services; improving real-time decision-making; and detecting and preventing fraud or security threats.
Despite those benefits, the increase in the number of platforms has introduced chaos. 86% of respondents struggle with integration across different data streaming platforms (29% call it a big issue, 57% an issue). The main reason is that each platform has its own governance model, access controls, encryption standards, auditing formats, schema enforcement, and retention policies. Integration becomes painful not just technically, but at the policy level too.
Asked which data streaming capabilities need the most improvement, respondents named integration with AI/ML platforms first, ahead of connectors to mainstream applications and security and governance. Even so, 80% rate their current AI/ML integration as "good," 9% "excellent," and 11% "average", a confidence that the operational findings later in this report complicate.
Respondents use a wide range of tools for data preparation, training, and inference.
Google Cloud Vertex AI66%
Microsoft Azure Machine Learning50%
Amazon SageMaker Data Wrangler42%
AWS Glue DataBrew38%
AWS SageMaker37%
That breadth of tooling helps explain the quality problems teams report when using streaming data for AI. The top five quality issues were:
- Inconsistent data formats or schemas
- Duplicate events causing skewed results
- Missing or incomplete data
- Low-quality data affecting model accuracy
- Irrelevant or outdated data
The findings expose the gap between AI ambitions and operational reality. Respondents rate their integration as good, yet name real-time processing and data quality as their biggest obstacles when scaling data infrastructure to support AI.
Data privacy and security concerns72%
High infrastructure costs59%
Lack of real-time processing58%
Data quality issues56%
Integration with existing systems49%
Talent and expertise shortages7%
Organizations are focusing on a handful of success indicators, such as data latency and freshness, while ignoring a key fact: they don't have unified governance across different platforms and technologies. They need uniform access control, schema enforcement, and retention policies, not one set for each team or environment.
Respondents report using a wide range of data lakes and warehouses, which compounds the problem:
- Data lakes: Amazon S3 and/or Lake Formation, Databricks Delta Lake, and Google Cloud Platform.
- Data warehouses: Google BigQuery, Amazon Redshift, and (joint) Azure Synapse Analytics / IBM Db2 Warehouse.
When moving data from streaming systems into lakes and warehouses, teams reach for a mix of approaches:
Custom pipelines (Spark or Flink)73%
Kafka Connect or similar tools69%
Managed services (Firehose, Snowpipe)50%
Micro-batching before loading49%
ELT / ETL tools (Fivetran, Airbyte)28%
The top three pain points during ingestion and transformation were time efficiency (collecting, connecting, and analyzing data in a centralized way), schema changes and rising data complexity, and parallel architectures that add complexity and require extra resources to manage.
The pain points above paint a picture of potential chaos for organizations using streaming to feed AI. The old adage of "garbage in, garbage out" is only amplified as the speed and volume of ingested data spiral. Examples of chaos include:
- Poor quality data. Pipelines ingest inconsistent, misformatted, and missing data at scale, reducing confidence and trust in data by teams and customers.
- Mystery costs. Topics and schemas are created, used, and abandoned without documentation or visibility, and unused assets accumulate unseen, wasting resources.
- Slow, overly manual workflows. Teams lack the systems and standards to let developers access data rapidly and safely, creating productivity bottlenecks.
Without the right governance in place, the business encounters increased time to market for new products, inaccurate outputs for real-time AI use cases like fraud detection, and degraded customer experiences.
The data streaming and AI gold rush is encouraging significant investment, but questions remain about how return on investment is measured. Platform metrics such as data latency matter, but they don't reveal ROI on their own.
| Operational success metrics (top 3) | Business success metrics (top 3) |
|---|
| End-to-end latency (ingestion to insight) | Revenue from real-time data services or products |
| Data freshness (event creation to availability) | Customer satisfaction / NPS from real-time experiences |
| System uptime and failure recovery times | Real-time decision-making accuracy or speed |
Operational KPIs help deliver the business metrics, but without attribution it's difficult to link streaming success to indicators such as revenue or customer churn.
Platform and security teams need to work together, unifying fragmented deployments into a single, cohesive environment. This is where Conduktor comes in. It's platform-agnostic, unifying governance across streaming Kafka environments, with federation built in to devolve responsibilities to developer and data teams while still letting platform teams implement guardrails. Conduktor is also bringing clarity and granularity to cost tracking, which simplifies cost attribution, ROI assessment, and forecasting.
The bottom line
- Adoption is outpacing governance. Streaming is the connective tissue between operational systems, analytics, and AI, but most organizations lack the governance, integration, and observability to make it work together.
- Fragmentation has a cost. Without a common control plane, the growth of streaming platforms turns into chaos, wasted investment, and lost market opportunities.
- A control plane closes the gap. Conduktor turns fragmented Kafka deployments into governed, AI-ready data platforms.
Turn fragmented streaming into governed, AI-ready data
See how Conduktor unifies governance, quality, and cost visibility across your streaming estate, with no application changes.
Book a demo