This is intended to be a high-level explanation of an assessment of the human contribution to risk, commonly known as the Human Reliability Assessment (HRA).
There are two distinct types of HRA:
- qualitative assessments that aim to identify potential human failures and optimize the factors that may influence human performance, and
- quantitative assessments, which, in addition, aim to estimate the likelihood of such failures occurring.
The results of quantitative HRAs can feed into traditional engineering risk assessment tools and methodologies, such as event and fault tree analysis or Human Error HAZOP.
There are difficulties in quantifying human failures (e.g., a lack of data regarding the factors that influence performance); however, there are significant benefits to the qualitative approach, and this type of HRA is described below. The organization should conduct qualitative analyses of human performance – identifying what can go wrong and then putting mitigation measures in place.
Example of a method to manage human failures
The following structure is well-established and has been applied in numerous industries, including chemical, nuclear, and rail. Other methods are available, but these tend to follow a similar structure to that described below. This approach is often referred to as a “human-HAZOP,” a useful term to help duty holders understand expectations.
Overview of key steps
- Step 1: consider main site hazards;
- Step 2: identify manual activities that affect these hazards;
- Step 3: outline the key steps in these activities;
- Step 4: identify potential human failures in these steps;
- Step 5: identify factors that make these failures more likely;
- Step 6: manage the failures using the hierarchy of control;
- Step 7: manage error recovery.
Step 1: consider main site hazards.
Consider the main hazards and risks on the site.
Step 2: identify manual activities that affect these hazards.
Identify activities in these risk areas with a human component. This step aims to identify human interactions with the system, which constitute significant sources of risk if human errors occur.
For example, there is more opportunity for human performance failures in chlorine bulk transfer than there is in a chlorine storage due to the number of manual
operations. Human interactions which will require further analysis are:
- those that have the potential to initiate an event sequence (e.g., inappropriate valve operation causing a loss of containment);
- those required to stop an incident sequence (such as activation of ESD systems) and;
- actions that may escalate an incident (e.g., inadequate maintenance of a fire control system).
Consider tasks such as maintenance, response to upsets/emergencies, and normal operations. It is important to note that a task may be a physical action, a check, a decision-making activity, a communications activity or an information-gathering activity. In other words, tasks may be physical or mental activities.
Step 3: outline the key steps in these activities
It is helpful to look at the activity in detail to identify failures. An understanding of the key steps in an activity can be obtained by talking to operators (preferably walking through the operation) and reviewing procedures, job aids, training materials, and the relevant risk assessment. This analysis of the task steps establishes what the person needs to do to carry out a task correctly. It will include a description of what is done, what information is needed (and where this comes from), and interactions with other people.
Step 4: identify potential human failures in these steps.
Identify potential human failures that may occur during these tasks – remembering that human failures may be unintentional or deliberate. Consider the guidewords below for the key steps of the activity. Key steps to consider would be those that could have adverse consequences should they be performed incorrectly.
A task may:
- Not be completed at all (e.g., non-communication);
- Be partially completed (e.g., too little or too short);
- Be completed at the wrong time (e.g., too early or too late);
- Be inappropriately completed (e.g., too much, too long, on the wrong object, in the wrong direction, too fast/slow);
or
- Task steps may be completed in the wrong order;
- The wrong task or procedure may be selected and completed;
Additionally, there may be:
- A deliberate deviation from a rule or procedure (a ‘procedural violation’).
A more detailed list of “error types,” similar to HAZOP guidewords, is provided at the end of this post. An operator may make the same failure on several occasions, which is known as dependency. For example, an operator may miscalibrate more than one instrument because they have miscalculated.
Step 5: identify factors that make these failures more likely
Where human failures are identified above, the next step is to identify the factors that make the failure more or less likely. Performance Influencing Factors (PIFs) are the characteristics of people, tasks, and organizations that influence human performance and, therefore, the likelihood of human failure. PIFs include time pressure, fatigue, design of controls/displays, and the quality of procedures. Evaluating and improving PIFs is the primary approach for maximizing human reliability and minimizing failures. PIFs will vary on a continuum from the best practicable to the worst possible. When all the PIFs relevant to a particular situation are optimal, then error likelihood will be minimized.
Some PIFs that should be considered when assessing an activity/task are outlined in the previous section on accident investigation. HSG48 also lists often-cited causes of human failures in accidents under the three headings of Job, Individual, and Organisation. These root causes of accidents are, in effect, the factors that can influence human performance and which should be reviewed in a human factors risk assessment. It is important to consider those factors under the control of management (such as resources, work planning, and training) as they can often influence a wide range of activities across the site.
Step 6: manage the failures using the hierarchy of control
To prevent the risks of human failure in a hazardous system, several aspects need to be considered.
- Can the hazard be removed?
- Can the human contribution be removed, e.g., by a more reliable automated system (bearing in mind the implications of introducing new human failures through maintenance, etc.)?
- Can the consequences of the human failure be prevented, e.g., by additional barriers in the system?
- Can human performance be assured by mechanical or electrical means? For example, the correct order of valve operation can be assured through physical key interlock systems or the sequential operation of switches on a control panel can be assured through programmable logic controllers. The actions of individuals should not be relied upon to control a major hazard.
- Can the Performance Influencing Factors be made more optimal (e.g. improve access to equipment, increase lighting, provide more time available for the task, improve supervision, revise procedures or address training needs)
Step 7: manage error recovery
Should it still be possible for failures to occur, improving error recovery and mitigation are the final risk reduction strategies? The objective is to ensure that, should an error occur, it can be identified and recovered from (either by the person who made the error or someone else, such as a supervisor) – i.e., making the system more “error tolerant.” A recovery process generally follows three phases:
- detection of the error,
- diagnosis of what went wrong and how, and
- correction of the problem
Detection of the error may include the use of alarms, displays, direct feedback from the system, and true supervisor monitoring/checking. There may be time constraints in recovering from certain errors in high-hazard industries, and it should be borne in mind that a limited time for response (particularly in an upset/emergency) is in itself a factor that increases the likelihood of error.
A Classification of Human Failures
This list of failures, akin to HAZOP guidewords, can be used in place of the simplified version in Step 4 of the method above.
Action Errors
A1 Operation too long/short
A2 Operation mistimed
A3 Operation in the wrong direction
A4 Operation too little / too much
A5 Operation too fast / too slow
A6 Misalign
A7 Right operation on the wrong object
A8 Wrong operation on the right object
A9 Operation omitted
A10 Operation incomplete
A11 Operation too early / late
Checking Errors
C1 Check omitted
C2 Check incomplete
C3 Right check on the wrong object
C4 Wrong check on the right object
C5 Check too early/late
Information Retrieval Errors
R1 Information not obtained
R2 Wrong information obtained
R3 Information retrieval incomplete
R4 Information incorrectly interpreted
Information Communication Errors
I1 Information not communicated
I2 Wrong information communicated
I3 Information communication incomplete
I4 Information communication unclear
Selection Errors
S1 Selection omitted
S2 Wrong selection made
Planning Errors
P1 Plan omitted
P2 Plan incorrect
Violations
V1 Routine Violations
V2 Situational Violations
V3 Deliberate actions
