Author: Stephan Eule

How AI Strengthens Crack Inspection

In a nutshell:

Inspection technology now provides more detailed pipeline data than ever before. Yet crack detection remains one of the most demanding areas of pipeline integrity, and interpreting complex signals requires experienced analysts.

As datasets grow larger and inspection signals become more complex, analyzing the data efficiently while maintaining consistency requires significant expertise, focus, and attention.

In this article, our subject matter expert Stephan Eule explores how artificial intelligence (AI) strengthens crack detection by helping analysts focus their expertise on the signals that matter most, enabling more consistent analysis and confident integrity decisions. 

Crack detection: A discipline built on evidence

In customer conversations, I often hear a simple question: if modern inspection tools collect more data than ever, why does crack analysis still require so much specialist expertise? The answer lies in the signal itself and in how carefully technology, field evidence and engineering judgement must be combined.

Crack detection is one of the most demanding areas of pipeline integrity. Stress corrosion cracking, fatigue cracking, and seam weld anomalies are planar, narrow, and often clustered, so they differ fundamentally from metal loss in both their failure mechanisms and the way they respond to inspection technologies, creating distinct inspection signals that require specialist interpretation. Reliable interpretation has required decades of tool development, structured field verification, and accumulated analyst expertise. Today, the industry performs this work to a high standard, and operators use the results every day to make critical integrity decisions.

Electromagnetic acoustic transducer (EMAT) technology is a leading in-line inspection (ILI) method for detecting cracks in gas pipelines. By generating guided ultrasonic waves directly in the pipe wall through electromagnetic coupling, it avoids the need for a liquid couplant and can operate under service conditions that conventional ultrasonic crack-detection tools cannot support. An ultrasonic technology such as EMAT, which creates acoustic reflections from a planar anomaly, is preferred over other technologies often used for metal loss detection, which rely on the volumetric properties of the targeted anomaly.

Over the past decade, EMAT hardware has improved substantially, with each tool generation adding more channels, denser sampling, and stronger signal quality. As a result, every inspection captures a richer view of the pipe than before. The value of that progress lies in turning this additional information into dependable integrity decisions, and this is where machine learning now makes its most direct contribution.

Why EMAT analysis requires specialist expertise

An analyst using wall-thickness ultrasonics measures geometry almost directly. Magnetic flux leakage (MFL) requires more interpretation, although decades of pull-test evidence have made the relationship between flux signals and metal loss geometry well established.

EMAT is different. Analysts do not derive depth directly from the signal. They interpret guided wave behavior, including reflections from features, transmitted amplitude, attenuation, and how these patterns vary across neighboring channels and along the pipe axis.

This is where expert judgement is essential, because several unrelated planar conditions can produce similar reflections. Crack colonies, milling features, surface roughness, general corrosion, and sensor liftoff over a dent may all affect the signal in comparable ways. Distinguishing between them depends on context, the line’s construction and operating history, and pattern recognition built through years of work with similar datasets. That expertise is a central asset of a crack detection service.

Depth and length sizing of the identified cracks raise the bar further. While detecting a crack and assigning it to a depth category is already challenging, estimating its actual depth requires a much higher level of precision and confidence. A continuous depth estimate requires stronger evidence because it must remain within a stated tolerance across features whose signal response varies with orientation, tightness and morphology. Achieving this level of reporting reflects the maturity of the discipline.

Defined qualification procedures, acceptance criteria, and structured review keep reported results at a high, auditable standard. Even so, some residual interpretive variation remains, because two qualified analysts may assess the same ambiguous signal slightly differently. These differences are small and conservative. Further reducing this variation is one of the clearest ways machine learning strengthens the service.

Better sensors provide richer evidence

Improved sensor resolution is a real step forward. Denser sampling makes smaller and weaker indications visible, broadening what an inspection can report and giving operators earlier insight into features developing in the pipeline. It also adds diagnostic detail to each signal, making the evidence for separating crack colonies from benign reflections stronger than before.

At the same time, richer data increases the number of decisions required during analysis. More candidate signals per kilometer create more judgements, all competing for the attention of the same expert analyst population.

Analyst expertise remains the most valuable input in the crack detection value chain. That expertise should be focused where it can change the outcome, rather than applied uniformly to every candidate signal. Machine learning makes this targeted use possible at the scale generated by modern inspection tools. As sensors improve, that capability becomes increasingly valuable. 

This image shows a portrait of Thomas Beuker, Head of Market and Service Line Strategy – Advanced Integrity Division.
The pursuit of quality, enabled by greater data density and volume, always goes hand in hand with the need for efficiency and consistency. Only an efficient and consistent analysis process can achieve the desired level of quality without becoming burdened by unnecessary and time-consuming conservatism.
Thomas Beuker, Head of Market and Service Line Strategy – Advanced Integrity Division

What AI actually means in this context

AI has become one of the most discussed topics in our industry, but it is also one of the most misunderstood. When I speak with customers about AI, I understand their caution. In a safety-critical environment, questions about trust, validation and human oversight are entirely reasonable. The following four misconceptions frequently create hesitation around the adoption of AI.

Misconception 1: AI means full automation. In practice, the model acts as an assistant, working alongside the analyst on the same dataset: drawing attention to signals that matter and helping clear routine cases. Screening mammography provides a useful analogy. A trained model can serve as a second reader while the radiologist remains accountable for the diagnosis. The aim is to reduce missed findings and reading time without transferring clinical responsibility to the model.

Misconception 2: AI must be a black box. A model trained on verified field data, validated on separate inspections and delivered with calibrated confidence can be treated as a measurement system with stated performance under defined conditions. Here, transparency comes from evidence, traceability and validation records – not from exposing every detail of the model architecture.

Misconception 3: A large general-purpose model can solve the problem. Such models have no useful prior understanding of guided-wave physics in carbon steel line pipe. Effective models for this domain must reflect the physics of the measurement and be trained on data from the relevant tool generations.

Misconception 4: A trained model is finished. Tool hardware evolves, pipeline populations change and analysis practices mature. Lifecycle management must therefore be part of the product, not an activity limited to the original development project. Treating models as maintained assets is essential to keeping performance stable over years of service.

Where AI creates the most value

The question is not whether AI can support crack assessment, but where it can make the greatest practical difference. The most valuable applications are those that help analysts work more efficiently, apply expertise more consistently and make better-informed decisions. Four applications stand out when value delivered is weighed against implementation risk.  

  1. Triage prioritizes candidate indications by their likelihood of being reportable crack-like features. The analyst still reviews the full population. The ranking determines which signals receive attention first and how efficiently the remaining cases are cleared.
  2. Discriminating benign signals offers the largest single efficiency gain. Coating disbondment, milling features, debris, and wall-thickness changes produce many signals that are not relevant to integrity. Pre-classifying these cases frees expert time for features with real integrity consequence.
  3. Consistency comes from applying the same criteria to every run, from the first to the four-hundredth. Used as a check on analyst output, the model highlights disagreements for review while the analyst retains judgement.
  4. Sizing support models increase confidence in reported depths. Trained on tightly controlled laboratory measurements, they provide depth estimates for individual EMAT shots and give data analysts a stronger basis for decision-making.

None of these applications removes human decision-making. They redirect human attention to the decisions where expertise matters most.

Models depend on the data behind them

Label quality and data coverage are decisive for machine-learning models used with ILI data. They separate models that perform reliably in the field from those that work only in controlled studies.

An analyst call is expert judgement, not ground truth. A model trained only on analyst calls will reproduce the current workflow, including its variation and blind spots. The most informative label is the verified field outcome: the excavation result, the in-ditch NDE measurement, and, where available, the recovered pipe section with crack morphologies reconstructed in three dimensions using X-ray computed tomography (XCT) scans.

Portrait of Stephan Eule, Senior AI Specialist.
The greatest value comes from connecting three elements: the raw inspection signal, the analyst interpretation based on that signal, and the verified field outcome that follows. Each can be captured separately, but maintaining a traceable link across decades of inspection campaigns and successive tool generations requires sustained investment. ROSEN has built this foundation through decades of inspection and verification work across a broad pipeline population. This enables models to be developed with declared performance, not merely supported by promising demonstrations.
Stephan Eule, Senior AI Specialist

Customers often ask how much data is needed. The answer is usually “more,” but the better question is whether the data covers the conditions the model must handle: pipe grades and vintages, wall-thickness ranges, seam types, geographical regions, diameters, and other relevant variables. This breadth allows a model to be qualified for a population it has genuinely encountered, and it can only be built through years of diverse field experience.

Building AI that earns trust

Deploying AI in a safety-critical environment requires more than a research prototype can provide. Five requirements define the engineering standard.

Specify performance in familiar industry terms. API 1163 defines how in-line inspection system performance is declared and validated, and any machine-learning component within such a system carries the same obligation. Probability of detection, probability of identification, sizing tolerance and false-call rate should be stated at a defined confidence level for a defined feature population.  

Account for correlation in validation data. Signals from the same pipeline, joint or inspection run are strongly related, so random train-test splits can overstate accuracy. Separating data by pipeline and campaign, and keeping the held-out test set untouched during development produces performance figures operators can trust.

Calibrate uncertainty. A model that recognizes its limits is more useful than one that answers every case with the same confidence. Conformal prediction can provide coverage guarantees under stated assumptions and allow a model to abstain. In a safety-critical service, routing difficult cases to a human is the correct behavior.

Monitor performance continuously. Monitoring should cover both input data and reported outcomes. Any drift away from the qualified population then becomes visible early, keeping performance under active management throughout deployment.

Separate responsibilities through governance. The team developing a model should be distinct from the team validating it. Version control must cover code, models, training data and validation evidence, and every reported call should remain traceable to a model version and a named human reviewer. Retraining requires the same change-control discipline as a hardware modification.

Trustworthy AI for pipeline integrity

In many fields, trustworthy AI is used as a broad label for robustness, uncertainty quantification, explainability, fairness, and governance. In pipeline integrity, it must be defined in more practical and testable terms.

A trustworthy model clearly states its capabilities, defines the population it is qualified to assess, quantifies confidence for each case, rejects inputs outside its qualified scope, and leaves an audit trail that enables an integrity engineer to defend the decision to a regulator.  

Reasoning across inspection technologies

Interpreting a single inspection technology depends on a manageable set of signal features that an experienced analyst can assess together. Interpreting multiple technologies in combination is far more complex.

Consider an EMAT indication where the geometry channel shows a dent, MFL identifies coincident metal loss, and three earlier inspections add their own history. Each observation changes the interpretation of the others. The dent can affect the EMAT response through sensor liftoff and local strain, while the combined signature of a dent with metal loss and cracking at its base differs from any of those features viewed in isolation.

Such combinations are numerous, individually rare, and rarely captured in rule books, making explicit rule sets impractical. This is exactly where machine learning is useful: a model can learn interaction patterns from examples instead of requiring every possible case to be defined in advance.

Operators that have inspected the same line with multiple technologies already hold much of this data. The earlier applications make established workflows faster and more consistent; this application enables an assessment that no analyst could perform manually at full-pipeline scale.

Defining the standards

The next challenge is not the technology itself. It is ensuring that AI is introduced in a way that maintains the high levels of trust, traceability and technical rigor expected in pipeline integrity. As adoption accelerates, the industry has an opportunity to establish common standards before practices become fragmented.

Several areas should be addressed now. Systems that learn from data need a shared qualification framework, including minimum requirements for validation datasets. Inspection reports should identify the model version used; retraining should have clearly defined requalification triggers; performance should be assessed for the combined analyst-and-model workflow rather than the model alone; and data provenance and label traceability should be standard requirements.

Reaching that point will require industry-wide collaboration. ROSEN intends to contribute actively by bringing field evidence, inspection expertise and practical deployment experience into the discussion.

The analyst’s evolving role

As AI models support data analysts, the analyst’s role becomes more focused – and more valuable. The model takes on routine screening, while the analyst resolves disagreements, interprets unusual signals, integrates pipeline history and operating context, performs the engineering assessment and makes the final call. These are the activities where deep expertise has the greatest impact, and AI creates more time to apply it.

Portrait of Stephan Eule, Senior AI Specialist.
Accountability remains unchanged. Reports are signed by qualified people, supported by qualified processes, using evidence that the model helped assemble.
Stephan Eule, Senior AI Specialist

The next decade of crack detection will be shaped by four connected advances: sensors that reveal more of the pipe, models that absorb routine screening effort, analysts who focus on decisions where expertise changes outcomes, and governance strong enough for operators to act with confidence. The opportunity is not to replace expert judgement, but to apply it more consistently, at greater scale and where it creates the most value.  

Ultimately, the greatest advances in crack detection will come not from AI alone, but from the combination of human expertise, trusted data, and well-governed technology.

Portrait of Stephan Eule, Senior AI Specialist.

Stephan Eule 

Senior AI Specialist

Stephan has worked at ROSEN for eight years, focusing on the research and development of AI products in the crack detection portfolio. Prior to joining ROSEN, Stephan led a research group at the Max Planck Institute for Dynamics and Self-Organization and held a visiting researcher position at Harvard University. He has over ten years of experience applying deep learning to sensory data and computer vision. Currently, Stephan identifies the right use cases for AI-assisted inspections and transforms them into practical solutions.

Contact me
This image shows a portrait of Thomas Beuker, Head of Market and Service Line Strategy – Advanced Integrity Division.

Thomas Beuker

Head of Market and Service Line Strategy – Advanced Integrity Division

Thomas is responsible for the service line strategy for ROSEN’s crack inspection services, as well as in-line inspection (ILI) services used to determine the material properties and stress and strain status of pipelines. He started his career at ROSEN in the Research & Development department in Germany and has more than 30 years of experience in developing and deploying pipeline inspection solutions. Thomas holds an MSc in Geophysics from the University of Muenster, Germany. Since 1997, he has represented ROSEN at the Pipeline Research Council International (PRCI).

Contact me
Close up of a hand holding a cell phone on which the facet newsletter can be seen.

Not yet registered to facets?

Register now if you would like to see more stories like this and receive the latest news and updates.
Read more