A growing biotech company wanted to answer a practical manufacturing question:

Why are some batches taking longer than others?

The information needed to investigate the question already existed.

Manufacturing information was in the Manufacturing Execution System (MES). Process conditions were captured by a historian. Laboratory results were in the Laboratory Information Management System (LIMS). Quality events were managed in the electronic Quality Management System (eQMS). Material and production information was available in the Enterprise Resource Planning (ERP) system.

The challenge was bringing the right information together so teams could understand what was happening across the operation.

This is a practical example of biotech data modernization and life sciences data engineering: making the data a company already has easier to connect, understand and use.

Case Study at a Glance

OrganizationGrowing biotech manufacturer
Business questionWhy are some batches taking longer to complete than others?
Data sourcesMES, LIMS, eQMS, ERP and historian
FocusData integration, quality, governance and traceability
GoalBuild a connected view of manufacturing performance
Future useManufacturing analytics and AI-ready data

Why Was the Question Difficult to Answer?

Each system provided a different part of the answer.

Business QuestionWhere the Data May Live
When did manufacturing steps start and finish?MES
What were the process conditions?Historian
When was laboratory testing completed?LIMS
Was the batch associated with a deviation or investigation?eQMS
What production order and materials were involved?ERP

For a manufacturing review, Operations might look at MES records while QC reviews laboratory results and Quality checks deviations.

The information may all be available, but someone still needs to make sure the records relate to the same batch.

For example, the same production activity might appear as:

MES: Batch B26-1047
LIMS: Lot 1047-DS
ERP: Production Order PO-48172

Those records may all belong together, but the relationship needs to be defined.

Similar differences can appear in product names, equipment identifiers, units, timestamps and manufacturing terminology.

This is where MES, LIMS, eQMS, ERP and historian data integration becomes useful. The goal is not simply to put information in one place. The data needs to work together.

Start With What the Company Needs to Know

Instead of beginning with a technology platform, the project focused on three questions:

  • How long is each batch taking?
  • Where are delays occurring?
  • Are manufacturing, laboratory or quality events contributing to those delays?

Those questions helped determine what information was actually needed.

For example, MES could provide manufacturing steps and timestamps. LIMS could provide laboratory testing and completion times. eQMS could identify deviations associated with a batch. Historian data could provide relevant process conditions.

There was no need to connect every field from every system.

The first goal was to connect the right data for the problem being solved.

Making the Data Work Together

The required information can be collected using APIs, database connections, ETL/ELT pipelines, file-based integrations or other methods that fit the company’s existing systems.

Once collected, the data needs a consistent structure and meaning.

Consider batch cycle time.

Operations might calculate it from the beginning to the end of manufacturing. Another team might include laboratory testing or the time until final disposition.

Both calculations can be useful, but they answer different questions.

The organization therefore needs to agree on:

  • what the metric means
  • which source data should be used
  • when the calculation starts and ends
  • how records from different systems relate
  • how units and timestamps should be handled

This is data governance in practical terms.

When the definitions are clear, teams can spend less time working out whose number is correct and more time understanding what the number means.

Making the Results Trustworthy

Connecting and organizing the data is only useful if people can trust the result.

Suppose a manufacturing report shows:

Batch Cycle Time: 126.4 hours

A reviewer should be able to understand what sits behind that number.

What started the clock? What stopped it? Which systems provided the information? Was anything transformed or calculated? Was all expected information received? Can the result be traced back to its source?

Data quality checks can also identify exceptions.

For example, if MES shows a completed batch but an expected QC result has not been received from LIMS, the missing information can be flagged rather than silently treated as a complete record.

For GxP data, the level of control should also reflect how the information will be used.

Information used for exploratory operational analysis may have different requirements from information supporting a GMP quality decision.

A practical GxP data governance approach can therefore consider data integrity, traceability, security and access, change control, intended use and the appropriate level of qualification or validation.

What Could This Mean in Practice?

The amount of effort required to review a batch varies considerably.

A short, straightforward batch with clean records is very different from a complex batch with a long batch record, multiple systems, deviations, investigations or missing information.

For this representative scenario, a broader cross-system review could reasonably involve:

ActivityEffort per Batch
Find and collect relevant records30–60 minutes
Review manufacturing and batch information1–2 hours
Review laboratory and process information30–60 minutes
Review deviations or other exceptions30 minutes–2+ hours
Reconcile and compile the information30–60 minutes
Estimated overall effortApproximately 4–7+ hours

For a review covering 20 batches, that represents approximately:

80–140+ hours

or roughly:

2–4+ weeks of one person’s working time

The purpose of a connected data foundation is not to automate all of those 80–140 hours.

Some of that time represents important human work: reviewing results, understanding exceptions, investigating issues and making decisions.

The opportunity is to reduce the repetitive work around that review.

Instead of repeatedly finding records, exporting data, matching identifiers and rebuilding the same analysis, the required information can be collected and organized in a reusable way.

That gives reviewers more time to focus on the questions that actually require their expertise.

Why did this happen? rather than: Where do I find the information?

These estimates are illustrative and are not measured client results. Actual effort varies based on batch record length, process complexity, number of systems, deviations and investigations, data availability and review scope.

What Technology Can Support Life Sciences Data Modernization?

There is no single technology stack required for pharma or biotech data modernization.

The right technology depends on the company’s existing environment, the systems and data being connected, business needs, internal capabilities and regulatory requirements.

Common technologies that may support this type of environment include:

NeedExample Technologies
Cloud infrastructureMicrosoft Azure, AWS
Data platformsDatabricks, Snowflake, Microsoft Fabric
Cloud data storageAzure Data Lake Storage, Amazon S3, Microsoft OneLake
Data integration and pipelinesAzure Data Factory, AWS Glue, Microsoft Fabric Data Factory, Fivetran
Data transformationdbt, Databricks, Microsoft Fabric
Reporting and visualizationMicrosoft Power BI, Tableau
Data catalog and governanceMicrosoft Purview, Databricks Unity Catalog

These are examples, not a required architecture.

A biotech already using Microsoft technology may make different choices from an organization built around AWS, Databricks or Snowflake.

When the data or resulting information supports a GxP-regulated process, technology selection is only one part of the solution.

The organization also needs to consider how the overall solution is designed, configured, controlled and used.

A cloud or data platform is not automatically GxP compliant, FDA 21 CFR Part 11 compliant or EU GMP Annex 11 compliant simply because it is used in the architecture.

The appropriate controls depend on the specific technology, configuration, intended use, data and regulatory requirements.

What Becomes Possible With Connected Data?

Once manufacturing, laboratory and quality information can be reliably used together, the company can start answering broader questions without rebuilding the underlying dataset each time.

For example:

  • Which manufacturing steps contribute most to batch cycle time?
  • Is QC turnaround affecting overall release timing?
  • Which process steps show the greatest variation?
  • Are certain pieces of equipment associated with more quality events?
  • Are particular deviation types becoming more common?
  • Do similar patterns appear across products or batches?

The same modern data foundation can also support more advanced analytics and AI.

For example, a biotech may eventually want to look for relationships between manufacturing conditions, laboratory results and quality events.

That is where an AI-ready data foundation in life sciences becomes valuable.

AI does not simply need more data. It needs data that can be connected, understood and trusted for the intended use.

A Practical Starting Point for Biotech Data Modernization

This example shows why data modernization does not have to begin with a large technology transformation.

The biotech already had useful data across its operation.

The opportunity was to make that information work better together.

By starting with a real manufacturing question, connecting the relevant systems, establishing common meaning and building appropriate data quality and traceability, the company can create a foundation that supports today’s analysis and future use cases.

Start with one useful question. Connect the data needed to answer it well. Then build from there.

Frequently Asked Questions

How do biotech companies connect MES and LIMS data?

MES and LIMS data can be connected using APIs, ETL/ELT pipelines, database connections or other integration methods. The important part is also defining how batches, lots, samples and other records relate between the systems so the information can be reliably analyzed together.

How can MES, LIMS, eQMS, ERP and historian data be used together?

These systems provide different parts of the operational picture. MES provides manufacturing execution information, historians capture process conditions, LIMS contains laboratory results, eQMS manages quality events and ERP provides production and material information. Connecting relevant data can support manufacturing analysis, investigations, trend analysis and other cross-system use cases.

What technology is used for life sciences data modernization?

Common options include Microsoft Azure, AWS, Databricks, Snowflake and Microsoft Fabric, along with technologies for data integration, transformation, governance and reporting. The right technology depends on the organization’s existing environment, data, business needs, intended use and regulatory requirements.

Does a pharma or biotech data platform need to be validated?

It depends on how the platform and resulting information will be used. A platform supporting general business analysis may have different requirements and not require validation. However, an analysis supporting a GxP decision would require validation. Intended use, data criticality and risk should help determine the appropriate controls and level of qualification or validation.

Is a cloud data platform automatically 21 CFR Part 11 or EU GMP Annex 11 compliant?

No. Selecting a particular cloud or data platform does not by itself make the overall solution compliant with FDA 21 CFR Part 11 or EU GMP Annex 11. Compliance depends on the complete solution, its intended use, configuration and applicable controls.

How does data modernization prepare life sciences data for AI?

Data modernization can connect information from systems such as MES, LIMS, eQMS and historians, establish consistent meaning and improve data quality and traceability. This creates a stronger AI-ready data foundation for analytics and AI applications.

Making Life Sciences Data More Useful

Assurea helps biotech and pharmaceutical organizations make better use of data across manufacturing, laboratory, quality and enterprise systems.

Our life sciences data modernization and engineering services include data strategy and architecture, data integration and engineering, cloud data platform modernization, data governance and quality, and analytics and AI-ready data foundations.

The goal is simple: Make the data you already have easier to connect, understand, trust and use.

Disclaimers

Case Study Note: This article presents a representative life sciences data modernization scenario based on common industry challenges and implementation approaches. Certain details, examples and effort estimates have been generalized or are illustrative for educational purposes. They are not presented as measured client results. Actual architecture, effort, controls and outcomes will vary based on the organization, systems, batch complexity, intended use and requirements.

Technology Disclaimer: Technology platforms and products referenced in this article are examples of technologies that may be used in life sciences data environments. Their inclusion does not indicate that they were used in the representative scenario or that any individual technology is inherently compliant with GxP, FDA 21 CFR Part 11 or EU GMP Annex 11. Suitability and required controls depend on the specific product or service, configuration, intended use, architecture and applicable regulatory requirements. Assurea is not affiliated with or endorsed by the technology vendors referenced in this article unless otherwise expressly stated.