When a workplace accident occurs, understanding what went wrong is only half the battle. The real value lies in preventing similar incidents from happening again. This requires systematic investigation methods that go beyond surface-level observations to uncover the true causes of accidents. Different analytical approaches serve different purposes, from evaluating equipment failures to understanding human behavior patterns.
Table of Contents
- Understanding accident investigation approaches
- Failure mode and effect analysis
- How FMEA works in practice
- Fault tree analysis
- Building and using fault trees
- Theory of human error
- Analyzing human factors systematically
- Cost effectiveness analysis
- Practical application in safety decisions
- Statistical method for pattern identification
- Identifying systemic weaknesses
- Critical incidence technique
- Building a safety knowledge base
- System safety method
- Proactive safety design and analysis
- Selecting and combining methods
Understanding accident investigation approaches
Accident investigation methods have evolved significantly over the decades. Modern approaches recognize that workplace accidents rarely result from a single cause. Instead, they typically involve multiple, interrelated factors spanning equipment, procedures, training, management systems, and human performance. Selecting the right investigation method depends on the complexity of the incident, the nature of the workplace, and the specific information needed to prevent recurrence.
Failure mode and effect analysis
Failure Mode and Effect Analysis (FMEA) is a structured approach that examines how individual components or processes can fail and evaluates the consequences of those failures on overall system performance. Originally developed by the U.S. military in the 1940s, FMEA has become a cornerstone method for identifying and prioritizing potential risks before they manifest as actual incidents.
The process involves identifying potential failure modes for each function, determining the consequences of each failure, and rating the severity of effects. Organizations assign severity ratings typically on a scale from one to ten, where one represents insignificant impact and ten indicates catastrophic failure that could cause injury or death. This systematic evaluation helps teams focus resources on the most critical failure points.
How FMEA works in practice
In a manufacturing setting, an FMEA team would examine each piece of equipment and process step to identify ways they could fail. For instance, a conveyor system might fail due to motor burnout, belt breakage, or sensor malfunction. Each failure mode is then analyzed for its potential effects on worker safety and production continuity. The team calculates a Risk Priority Number by multiplying severity, occurrence likelihood, and detection difficulty ratings.
FMEA excels at analyzing complex, interrelated systems where one component’s failure can cascade through the entire operation. However, the method has limitations. It typically analyzes only one failure at a time, which means it may miss interactions between multiple simultaneous failures. Additionally, the analysis can become time-consuming for very large systems with numerous components.
Fault tree analysis
Fault Tree Analysis (FTA) takes a fundamentally different approach by starting with an undesired outcome and working backward to identify all possible contributing events. This top-down, deductive method creates a visual tree structure that maps the logical relationships between various failures and the final incident.
The technique originated in 1962 at Bell Laboratories for evaluating the Minuteman missile system and has since found widespread application across industries including aerospace, nuclear power, chemical processing, and automotive manufacturing. FTA uses Boolean logic gates to show how different events combine to produce system failures.
Building and using fault trees
A fault tree begins with a single top event representing the accident or undesired outcome being investigated. Below this, investigators identify intermediate events that could lead to the top event, continuing to break down causes until reaching basic events that cannot be further subdivided. Logic gates connect these events to show whether failures must occur simultaneously (AND gates) or whether any single failure is sufficient (OR gates).
The visual nature of fault trees makes them powerful communication tools. Even individuals unfamiliar with the methodology can follow the logical progression from basic causes to the final accident. FTA also enables quantitative analysis by assigning probabilities to basic events and calculating the likelihood of the top event occurring.
Despite its strengths, FTA can become complex for large systems, requiring specialized expertise and sophisticated calculations. The method works best when analyzing specific, well-defined failure scenarios rather than exploring all possible system vulnerabilities.
Theory of human error
Human error contributes to the vast majority of workplace accidents, making it essential to understand why people make mistakes and how organizational factors influence error likelihood. Analysis of accidents shows that human failure contributes to almost all incidents, including major disasters like Chernobyl, Piper Alpha, and the Texas City refinery explosion.
Human error theory distinguishes between different types of failures. Slips and lapses are unintended actions that occur during familiar tasks, such as pressing the wrong button or forgetting a procedural step. These happen even to well-trained personnel and cannot be eliminated through additional training alone. Mistakes, on the other hand, involve incorrect decisions where someone does the wrong thing while believing it to be right, often due to inadequate training or unclear procedures.
Analyzing human factors systematically
Rather than stopping at “operator error,” effective human error analysis identifies the underlying performance influencing factors that make mistakes more likely. These include poor equipment design, time pressure, excessive workload, inadequate training, unclear communication, and organizational culture issues. The Human Factors Analysis Classification System provides a structured framework for examining how supervisory factors and organizational decisions create conditions that increase error probability.
This analytical technique can be combined with other methods like FMEA or FTA to provide a more comprehensive picture of accident causation. For example, a fault tree might identify a valve closure as a basic event, while human error analysis would explore why the operator closed the wrong valve considering factors like confusing labeling, poor lighting, or distraction from simultaneous alarms.
Cost effectiveness analysis
Cost-benefit analysis enables contractors and organizations to assess the true cost of accident prevention measures against the financial benefits of reduced injuries, improved productivity, and enhanced system effectiveness. This economic approach helps decision-makers select safety improvements that provide the greatest return on investment.
The analysis compares the costs of implementing safety measures, including equipment purchases, training programs, and procedural changes, against the benefits derived from accident prevention. Benefits include avoided medical expenses, reduced insurance premiums, prevention of production disruptions, and preservation of organizational reputation. The method also considers indirect costs such as investigation time, replacement worker training, and potential regulatory penalties.
Practical application in safety decisions
When evaluating whether to install additional machine guarding, for instance, an organization would calculate installation and maintenance costs and compare these against the expected reduction in injury rates and associated expenses. The analysis might reveal that some safety measures provide substantial net benefits while others offer limited returns relative to their cost.
This approach proves particularly valuable when resources are limited and organizations must prioritize among multiple potential safety improvements. However, it requires careful consideration of all costs and benefits, including those that are difficult to quantify such as worker morale, safety culture impacts, and long-term reputational effects.
Statistical method for pattern identification
The statistical method represents a conventional approach that groups accident causes into predefined categories to identify patterns and trends across multiple incidents. Common classification frameworks include the 4 M’s (Man, Machine, Material, Management) and the 3 E’s (Engineering, Enforcement, Education).
By categorizing numerous accidents according to these frameworks, organizations can identify which types of causes occur most frequently and which problem areas require the most attention. Statistical analysis might reveal that a disproportionate number of accidents involve inadequate lockout-tagout procedures, pointing to the need for improved training and enforcement in this area.
Identifying systemic weaknesses
This method excels at revealing organizational weaknesses that might not be apparent when investigating individual incidents in isolation. When data shows that slips and falls consistently occur in certain departments or during particular shifts, it suggests systemic issues with housekeeping practices, supervision, or work scheduling that require management-level intervention.
The statistical approach works best as a complement to other investigation methods rather than as a standalone technique. While it effectively identifies patterns, it provides less detail about the specific causal mechanisms in individual accidents. Organizations typically use statistical analysis to guide where to focus more detailed investigations using methods like FMEA or FTA.
Critical incidence technique
The Critical Incidence Technique involves systematically collecting direct observations of human behavior through structured interviews with workers who have experienced or witnessed errors and near-misses. Rather than waiting for actual accidents to occur, this proactive method identifies hazardous situations and error-prone conditions before they result in injuries.
Investigators conduct interviews with a random sample of employees, asking them to describe specific incidents where errors occurred or could have occurred, regardless of whether harm resulted. These incidents are then categorized to reveal common patterns of mechanical failures and human errors with accident potential. The technique emphasizes specific, observable incidents rather than general opinions about safety conditions.
Building a safety knowledge base
The strength of this approach lies in its ability to capture valuable safety intelligence from frontline workers who directly interact with hazards daily. Workers often observe dangerous conditions or near-misses that never get formally reported but represent significant accident potential. By systematically collecting and categorizing these observations, organizations build a comprehensive understanding of vulnerabilities throughout their operations.
The method requires careful attention to creating a climate of trust where workers feel comfortable sharing information about errors without fear of punishment. When implemented effectively, it provides early warning of problems that other methods might only identify after an actual accident occurs.
System safety method
System safety methodology takes a holistic view that examines the interrelationships between all elements that contribute to accidents, including tools, equipment, materials, people, procedures, documents, and organizational structures. Rather than focusing narrowly on immediate causes, this approach considers how different system components interact and how weaknesses in multiple areas can combine to create accident conditions.
The method proves particularly valuable for analyzing complex socio-technical systems where accidents emerge from the interaction of technical, human, and organizational factors. Systems-based approaches like STAMP (Systems-Theoretic Accident Model and Processes) examine why existing controls failed to prevent or detect hazards and how safety constraints were violated.
Proactive safety design and analysis
One distinctive feature of system safety methods is their applicability before accidents occur. During the design phase of new processes or facilities, system safety analysis can identify potential vulnerabilities and guide the implementation of appropriate safeguards. This proactive capability makes the method particularly valuable for preventing accidents rather than merely learning from them after they happen.
The holistic perspective helps organizations understand that improving safety requires more than fixing individual components or retraining specific workers. It often necessitates changes to organizational structures, communication systems, decision-making processes, and the fundamental design of work systems. While this comprehensive approach provides deeper insights, it also requires significant expertise and resources to implement effectively.
Selecting and combining methods
No single investigation method suits every situation. Simple incidents may require only basic root cause analysis, while complex accidents involving multiple systems and organizational factors benefit from comprehensive approaches combining several methods. Organizations should develop investigation capabilities appropriate to their risk profile and systematically apply methods matched to incident complexity.
Many successful safety programs integrate multiple analytical approaches. An organization might use statistical methods to identify problem patterns, apply FMEA to evaluate equipment-related risks, incorporate human error analysis for procedural failures, and employ system safety thinking for major incident investigations. The key is selecting methods that provide the depth of analysis needed while remaining practical to implement with available expertise and resources.
What do you think? Which investigation methods does your organization currently use, and have you identified gaps where additional analytical approaches could strengthen your accident prevention efforts? How might combining multiple investigation techniques provide deeper insights than relying on a single method?
References
- https://en.wikipedia.org/wiki/Failure_mode_and_effects_analysis
- https://quality-one.com/fmea/
- https://www.ibm.com/think/topics/fault-tree-analysis
- https://fiixsoftware.com/glossary/fault-tree-analysis/
- https://www.hse.gov.uk/humanfactors/topics/humanfail.htm
- https://ascelibrary.org/doi/10.1061/%28ASCE%29CO.1943-7862.0000496
- https://en.wikipedia.org/wiki/Critical_incident_technique
- https://www.mdpi.com/2071-1050/14/10/5869
- https://www.sciencedirect.com/science/article/abs/pii/S0950423014001193
Leave a Reply