Wednesday, November 24, 2021

Diagnostic Performance Metrics for PHM Algorithms

Carl Byington developed and patented diagnostic and prognostics technologies for Impact Technologies, Sikorsky Aircraft, and Lockheed Martin. Carl Byington became an expert in prognostics and health management (PHM) technologies and next-generation condition-based maintenance (CBM) solutions. He currently consults in these technical areas with his PHM Design company, located in Georgia. In this latest blog series he discusses how to verify and validate diagnostic algorithms for a specific application. See here for more about Carl Stewart Byington

Algorithm Concept

Diagnostic algorithms are typically qualified on specific types of faults with limited test data and engineering judgment of intended applicability. Mechanical system ailments in gearboxes for instance may vary from root cause conditions such as shaft misalignment; to fatigue events such as spalled bearings; to slower wear processes related to scuffing wear on gears. The accuracy of the fault detection and diagnostic processes for such a wide array of problems will not only depend on the algorithm’s sensitivity to signal- to-noise ratio but also load level, failure type and flight condition. Diagnostic algorithms that are sensitive to faulted conditions yet relatively insensitive to confounding conditions are desirable for a broader range of application in equipment health monitoring. Such generalized algorithms, though, may be less sensitive to early fault detection. In order to assess the risk associated with using certain diagnostic algorithms, qualify them for a range of use, determine desirable thresholds to produce known false alarm rates, and develop fusion approaches, we need to evaluate the detection performance and diagnostic accuracy using established performance metrics. Here are some specific metrics we can use for each step.

Decision Matrix Construct

The following Decision Matrix defines the cases used to evaluate fault detection. It is based on hypothesis testing methodology and represents the possible fault-detection combinations that may occur.

 Detection Decision Matrix


Outcome

Fault (F1)

No Fault (F0)

Total

Positive (D1)

(detected)

a

Number of defected faults

b

Number of false alarms

a+b

Total number of alarms

Negative (D0)

(not detected)

c

Number of missed faults

d

Number of correct rejections

c+d

Total number of non-alarms

 

a+c

Total number of faults

b+d

Total number of fault-free cases

a+b+c+d

Total number of cases

From this matrix, the detection metrics can readily be computed. The Probability of detection given a fault (a.k.a. sensitivity) assess the detected faults over all potential fault cases:

                                        (1)

The probability of false alarm (POFA) considers the proportion of all fault-free cases that trigger a fault detection alarm.

                                       (2)

The Accuracy is used to measure the effectiveness of the algorithm in correctly distinguishing between a fault-present and fault-free condition. The metric uses all available data for analysis (both fault and no fault):

Accuracy=         (3)

Diagnostic metrics are used to evaluate classification algorithms, typically consider multiple fault cases, and are based upon the confusion matrix concept. The matrix, illustrates the results of classifying data into several categories. A confusion matrix shows actual defects (as headings across the top of the table) and how they were classified (as headings down the first column). The shaded diagonal represents the number of correct classifications for each fault in the column and subsequent numbers in the columns represent incorrect classifications. Ideally, the numbers along the diagonal should dominate. The confusion matrix can be constructed using percentages or actual cases witnessed.

The probability of isolation (Fault Isolation Rate - FIR) is the percentage of all component failures that the classifier is able to unambiguously isolate. It is calculated using:

                                                (4)

                                                               (5)

              Ai = the number of detected faults in component i that the monitor is able to isolate                             unambiguously as due to any failure mode (Numbers on the diagonal of the Confusion Matrix).

              Ci = the number of detected faults in component i that the monitor is unable to                                     isolate unambiguously as due to any failure mode (Number off the diagonal of the Confusion Matrix).

An alternative metric that is also useful is the Kappa Coefficient, which represents how well an algorithm is able to correctly classify a fault with a correction for chance agreement.

         (6)

where:

N(obs in agreement) = sum of diagonals in matrix

N(exp in agreement) = sum{[sum of row)/N]*(sum of column)} for diagonals

N(total) = total number of observations

The metrics presented here form the basis for analysis of current detection and diagnostic algorithm effectiveness. The issue that confronts the PHM designer and researcher now is to generate realistic estimates of these metrics given the currently limited baseline data and faulted data available. The basis for an effective statistical analysis, given these limitations, will be discussed in a separate entry.

Additional Resources

A full review of detection and diagnostic metrics is summarized in a paper by Carl Byington, et al. It is available here:

http://www.humsconference.com.au/Papers2003/HUMSp404.pdf

Carl Byington’s related publications can be found at:

https://www.researchgate.net/profile/Carl-Byington

Carl Byington may be contacted for specific consulting engagements at:

https://phmdesign.com/contact-us/

Wednesday, September 8, 2021

CBM Program for US Army Aircraft


A former Rochester, NY consultant, Carl Byington emphasizes data and analytics-based approaches at PHM Design, LLC in Atlanta, GA. He works with clients to implement predictive analytics to maximize operational and maintenance efficiency. Well published in his field, Carl Byington presented the condition-based maintenance (CBM) of the United States Army aircraft systems at an American Helicopter Society (AHS) specialists’ meeting.

A CBM program involves moving from part replacements performed at defined intervals to maintenance performed upon “evidence of need". With the context of the Army’s CBM+ plan, this requires a move away from “time before overhaul” (TBO) protocol that traditionally defines schedules for military vehicle component maintenance. It also suggests an ability to move away from dedicated inspection and test flight maintenance events.

According to the paper, major obstacles in CBM adoption include the limited ability of digital source collectors (DSC) and health and usage monitoring systems (HUMS) to diagnose component faults early. Issues include condition indicators and sensors’ sensitivity to operating/environmental conditions, observability of specific failure modes, and inherent signal to noise ratio. Fleet-wide diagnostic thresholds to action are also difficult to implement, with uncertainties arising in detecting existing and progressive damage or wear among individual aircraft variations of use. Implementing prognostics on faulty or degrading components is inherently a significant endeavor, with validating and verifying such systems a remaining challenge.

The paper presents tools that can improve data monitoring, boost diagnostics, and enable prognostics to be implemented in ways that support a better transition to CBM. The realizable benefits are substantial. Detection of faults in their early stages provides an opportunity to order parts, schedule personnel, shutdown the equipment before serious damage occurs, and minimize the disruption to production and missions. Furthermore, insight gained from better diagnostics reduces uncertainty regarding the “health” of critical internal drivetrain components and allows for the safe reduction of some preventive maintenance and inspections. In other words, maintenance is performed only when necessary. As we extend such capability into predictive prognostics technology, often enabled now by machine learning (ML) and artificial intelligence (AI) techniques, we can realize even greater maintenance and logistics benefits.

Monday, June 21, 2021

Advanced Wind Turbine Health Management

 

Experienced entrepreneur and engineer Carl Byington worked for Impact Technologies, Sikorsky Aircraft, and Lockheed Martin for many of his professional years. In these roles, Carl Byington became an expert in prognostics and health management (PHM) technologies and next-generation condition-based maintenance plus (CBM+) solutions. He currently consults in these technical areas at his PHM Design company, located in the Atlanta, GA area.

 

Rotorcraft and wind turbines are both susceptible to critical drive train failure modes involving bearings and gears. The helicopter community has developed sophisticated health and usage monitoring systems (HUMS) to track damage accumulation on critical components and detect incipient faults. Health and usage monitoring systems typically consist of a variety of onboard sensors, data acquisition systems, and signal processing and analysis algorithms.  The acquired data may be processed onboard the rotorcraft or on a ground station (or a combination of both) providing the means to measure against defined criteria and generate instructions for the maintenance staff and/or flight crew for intervention. Application of HUMS technology to wind turbine drive trains has the potential to significantly reduce maintenance costs and increase turbine availability by enabling condition-based maintenance (CBM).

 

The wind turbine industry has experienced an array of drivetrain failures including spalled bearings and fractured gear teeth. Some of these failure modes are largely attributable to unexpected and/or excessive loading conditions. Such drivetrain failures often entail expensive repairs and can have catastrophic consequences for the turbine. The helicopter community faces many of the same challenges from uncertain loads and dire consequences associated with drivetrain failures. Thee helicopter community addressed these risks through the use of drivetrain health and usage monitoring systems (HUMS). Such HUMS oil debris and advanced gear/bearing vibration technologies can be used to help the wind turbine industry.

 

Information from HUMS enables implementation of a new maintenance paradigm that can improve reliability, availability, and maintainability of wind turbines while simultaneously reducing maintenance costs. Under the simplest maintenance paradigm, reactive maintenance, equipment is allowed to run until it fails without maintenance intervention. Reactive maintenance yields low reliability and high costs due to missed opportunities to detect and repair faults, secondary damage (progression of failure to a severe state), large logistics footprint (parts and labor) to cope with unexpected failures, and lost production while equipment awaits repair.

 

Preventive maintenance can improve equipment reliability (number of failures) by periodically overhauling equipment before it wears out. To avoid unexpected failures in equipment that has an uncertain service life, preventive maintenance must be performed well in advance of the mean time to failure. Consequently, preventive maintenance achieves high reliability at the cost of performing premature maintenance which includes frequent interruption of production for planned maintenance, high labor costs, and high parts usage.

 

Condition-based maintenance (CBM), on the other hand, enables high equipment reliability and low maintenance costs by eliminating the need for unnecessary overhaul activities while simultaneously allowing repairs to be performed on a planned basis. Condition monitoring provides insight into the “health” of individual pieces of equipment so that maintenance decisions can be made on a case-by-case basis rather than on fleet wide averages. Detection of faults in their early stages provides an opportunity to order parts, schedule personnel, shutdown the equipment before serious damage occurs, and minimize the disruption of production. Furthermore, insight gained from equipment condition monitoring reduces uncertainty regarding the “health” of critical internal drivetrain components and therefore eliminates the need for preventive maintenance. In other words, maintenance is performed only when necessary. 


The value proposition for wind turbines and specific technologies involved is summarized in a paper by Carl Byington, et al. It is available for download here:

 

https://www.researchgate.net/publication/253354657_Advanced_Vibration_Monitoring_for_Wind_Turbine_Health_Management

 

Carl Byington may be contacted for specific consulting engagements at:

https://phmdesign.com/contact-us/