Unlocking the Value of Data in the Age of AI: Insights from Henrique Lemes
For decades, companies across various sectors have recognized the transformative potential of the data at their disposal. It serves as a foundational element in enhancing user experiences and shaping strategic plans based on empirical evidence. As Artificial Intelligence (AI) technology becomes more accessible and practical for real-world applications, the value derived from available data has surged dramatically. However, harnessing this potential demands meticulous effort in data collection, curation, and preprocessing, alongside careful consideration of governance, privacy, anonymization, regulatory compliance, and security.
The Complexity of Data
In a recent discussion with Henrique Lemes, the Americas Data Platform Leader at IBM, we delved into the multifaceted challenges enterprises confront when adopting practical AI across varied use cases. Lemes emphasized that labeling all enterprise information simply as "data" fails to capture its complexity. Today’s enterprises navigate a fragmented landscape consisting of diverse data types, each with varying quality—especially when distinguishing between structured and unstructured sources.
Structured Data refers to information neatly organized in a standardized format, making it readily searchable and efficiently processable by software systems. Examples include databases, spreadsheets, and forms. Conversely, Unstructured Data lacks a predefined format or organization, which complicates its processing and analysis. This category encompasses a broad range of formats, including emails, social media posts, videos, images, documents, and audio files. While often overlooked, unstructured data harbors invaluable insights that, when effectively managed through advanced analytics and AI, can fuel innovation and guide strategic business decisions.
Henrique pointed out that a staggering statistic highlights the issue: “Currently, less than 1% of enterprise data is utilized by generative AI, and over 90% of that data is unstructured, which directly affects trust and quality.”
Trust and Data Utilization
The aspect of trust regarding data is crucial for organizational decision-making. Decision-makers must be confident that the information available to them is complete, reliable, and ethically sourced. However, studies show that fewer than half of the data available to businesses is employed for AI initiatives, with unstructured data frequently sidelined due to its inherent complexity in processing and compliance—particularly at scale.
To enable better data-driven decision-making, organizations need to transition from a trickle to a flood of accessible information. An effective solution proposed by Henrique is Automated Ingestion, which ensures that vast amounts of data can be captured and processed without overwhelming existing systems. Nonetheless, the importance of enforcing governance rules and data policies remains paramount, applicable to both structured and unstructured data alike.
Key Processes for Leveraging Data Value
Henrique outlined three crucial processes essential for enterprises aiming to extract maximum value from their data:
-
Ingestion at Scale: Automation is vital in streamlining this initial stage, allowing organizations to efficiently gather vast amounts of data without manual intervention.
-
Curation and Data Governance: Ensuring that data is not only collected but also refined and organized according to governance policies is critical for maintaining its quality and consistency.
- Enabling Generative AI: Once the data is properly ingested and governed, making it accessible for generative AI applications can unlock over 40% ROI compared to conventional Retrieval-Augmented Generation (RAG) use cases.
IBM offers a comprehensive strategy rooted in a deep understanding of enterprises’ AI journeys. Combining advanced software solutions and domain expertise, IBM helps organizations transform both structured and unstructured data into AI-ready assets while adhering to governance and compliance frameworks.
Simplifying Complexity through Integration
“We bring together the people, processes, and tools. It’s not inherently simple, but we simplify it by aligning all the essential resources,” Henrique explained. As companies expand and transform, the diversity and volume of their data continue to rise. To keep pace, AI data ingestion processes must remain both scalable and adaptable.
“**[Companies] encounter difficulties when scaling because their AI solutions were initially built for specific tasks. When they attempt to broaden their scope, they often aren’t ready, and the data pipelines grow more complex. This drives an increased demand for effective data governance,” he stated.
IBM’s methodology centers on a thorough understanding of each client’s unique AI trajectory, crafting a clear roadmap for achieving ROI through effective AI implementations. Prioritizing data accuracy—regardless of whether it is structured or unstructured—coupled with an emphasis on ingestion, lineage, governance, compliance, and observability, forms the cornerstone of their approach. These capabilities empower clients to scale across multiple use cases and fully leverage the intrinsic value of their data.
The Journey Toward Successful AI Implementation
Time is a critical element in implementing robust technological solutions. Establishing the right processes, selecting suitable tools, and foresee how data solutions may need to evolve takes significant investment and strategic planning. IBM equips enterprises with a broad array of options and tools to facilitate AI workloads, tailored to even the most regulated industries at any scale. With an illustrious clientele comprising international banks, finance houses, and global multinationals, IBM stands out as a preeminent player in this domain.
For organizations eager to uncover the potential of their data and enable robust data pipelines for AI that drive business results, along with a swift, substantial ROI, exploring enhanced solutions can be a pivotal step.
To find out more about enabling data pipelines for AI that serve business needs, head over to this page.
Inspired by: Source

