Performing a Human Reliability Assessment (HRA)

This information is intended to assist in assessing the human contribution to risk, commonly known as Human Reliability Assessment (HRA).  There are two (2) distinct types of HRA:

  1. qualitative assessments that aim to identify potential human failures and optimise the factors that may influence human performance, and
  2. quantitative assessments which, in addition, aim to estimate the likelihood of such failures occurring. The results of quantitative HRAs can feed into traditional engineering risk assessment tools and methodologies, such as event and fault tree analysis.

Quantifying human failures can be difficult (e.g., due to a lack of data regarding the factors that influence performance); however, the qualitative approach has significant benefits, and I will discuss this type of HRA.

At the end of the HRA, it is expected that the client is left with a human failure risk assessment.

Method to manage human failures

The following structure is well-established and applied in numerous industries, including chemical, nuclear, and rail. Other methods are available, but these tend to follow a similar structure. This approach is often referred to as a “human-HAZOP.” 

 

Overview of key steps

Step 1: consider main site hazards;

Step 2: identify manual activities that affect these hazards;

Step 3: outline the key steps in these activities;

Step 4: identify potential human failures in these steps;

Step 5: identify factors that make these failures more likely;

Step 6: manage the failures using hierarchy of control;

Step 7: manage error recovery.

 

Step 1: consider main site hazards
Consider the main hazards and risks on the site.

 

Step 2: identify manual activities that affect these hazards

Identify activities in these risk areas with a human component. This step aims to identify human interactions with the system which constitute significant sources of risk if human errors occur. For example, there is more opportunity for human performance failures in chlorine bulk transfer than in chlorine storage due to the number of manual operations.

Human interactions that will require further analysis are:

  • those that have the potential to initiate an event sequence (e.g., inappropriate valve operation causing a loss of containment);
  • those required to stop an incident sequence (such as activation of ESD systems) and;
  • actions that may escalate an incident (e.g., inadequate maintenance of a fire control system)

Consider tasks such as

  • maintenance,
  • response to upsets/emergencies,
  • normal operations

It is important to note that a task may be a physical action, a check, a decision-making activity, a communications activity, or an information-gathering activity. In other words, tasks may be physical or mental activities.

 

Step 3: outline the key steps in these activities

To identify failures, it is helpful to look at the activity in detail. An understanding of the critical steps in an activity can be obtained through talking to operators (preferably walking through the operation) and reviewing procedures, job aids, and training materials, as well as reviewing the relevant risk assessment. This analysis of the task steps establishes what the person needs to do to carry out a task correctly. It will include a description of what is done, what information is needed (and where this comes from) and interactions with others.

 

Step 4: identify potential human failures in these steps

Identify potential human failures that may occur during these tasks – remembering that human failures may be unintentional or deliberate. Consider the guide words below for the key steps of the activity. Key steps to consider would be those that could have adverse consequences should they be performed incorrectly.

A task may:

⇒Not be completed at all (e.g., non-communication);

⇒Be partially completed (e.g., too little or too short);

⇒Be completed at the wrong time (e.g., too early or too late);

⇒Be inappropriately completed (e.g., too much, too long, on the wrong object, in the wrong direction, too fast/slow);
or,

⇒Task steps may be completed in the wrong order;

⇒The wrong task or procedure may be selected and completed;

Additionally, there may be:

⇒A deliberate deviation from a rule or procedure (a ‘procedural violation’).

 

A more detailed list of “error types,” similar to HAZOP guidewords, is provided below. Note that an operator may make the same failure several times, which is known as dependency. For example, an operator may miscalibrate multiple instruments because they have made a miscalculation.

 

Step 5: identify factors that make these failures more likely

Where human failures are identified above, the next step is to identify the factors that make the failure more or less likely. Performance Influencing Factors (PIFs) are the characteristics of people, tasks, and organizations that influence human performance and, therefore, the likelihood of human failure.  PIFs include

  • time pressure,
  • fatigue,
  • design of controls/displays and
  • the quality of procedures

Evaluating and improving PIFs is the primary approach for maximizing human reliability and minimizing failures. PIFs will vary on a continuum from the best practicable to worst possible. When all the PIFs relevant to a particular situation are optimal, error likelihood will be minimized.

We can list often-cited causes of human failures in accidents under the three (3) headings:

  1. Job,
  2. Individual, and
  3. Organization

“Root causes” of accidents are, in effect, the factors that can influence human performance and which should be reviewed in a human factors risk assessment. It is essential to consider those factors under the control of management (such as resources, work planning, and training) as they can often influence a wide range of activities across the site.

 

Step 6: manage the failures using the hierarchy of control

Several aspects need to be considered to prevent the risks of human failure in a hazardous system.

  1. Can the hazard be removed?
  2. Can the human contribution be removed by a more reliable automated system, bearing in mind the implications of introducing new human failures through
    maintenance, etc.?
  3. Can additional barriers in the system prevent the consequences of human failure?
  4. Can human performance be assured by mechanical or electrical means? For example, the correct order of valve operation can be assured through physical key interlock systems, or the sequential operation of switches on a control panel can be assured through programmable logic controllers. The actions of individuals should not be relied upon to control a major hazard.
  5. Can the Performance Influencing Factors be made more optimal by:
    1. improving access to equipment,
    2. increasing lighting,
    3. providing more time for the task,
    4. improving supervision,
    5. revising procedures or
    6. addressing training needs

Step 7: manage error recovery

Should it still be possible for failures to occur, improving error recovery and mitigation are the final risk reduction strategies. The objective is to ensure that, should an error occur, it can be identified and recovered from either the person who made the error or someone else, such as a supervisor, making the system more “error tolerant.” 

A recovery process generally follows three (3) phases:

  1. detection of the error,
  2. diagnosis of what went wrong and how, and
  3. correction of the problem

Detection of the error may include using alarms, displays, direct feedback from the system, and true supervisor monitoring/checking.

There may be time constraints in recovering from specific errors in high-hazard industries, and it should be borne in mind that a limited time for response (particularly in an upset/emergency) is in itself a factor that increases the likelihood of error.

 

It is recommended that the following documents should also be requested:

  • risk assessment documents outlining the main hazards on site;
  • any analyses or documentation referring to safety-critical tasks, roles, or responsibilities.A Classification of Human Failures

This list of failures, akin to HAZOP guidewords, can be used in place of the simplified version in Step 4 of the method above.

Action Errors

 A1  Operation too long / short
 A2  Operation mistimed
 A3  Operation in wrong direction
 A4  Operation too little / too much
 A5  Operation too fast / too slow
 A6  Misalign
 A7  Right operation on wrong object
 A8  Wrong operation on right object
A9 Operation omitted
A10 Operation incomplete
A11 Operation too early / late

Checking Errors

 C1  Check omitted
 C2  Check incomplete
 C3  Right check on wrong object
 C4  Wrong check on right object
 C5  Check too early / late

Information Retrieval Errors

 R1  Information not obtained
 R2  Wrong information obtained
 R3  Information retrieval incomplete
 R4  Information incorrectly interpreted

Information Communication Errors

 I1  Information not communicated
 I2  Wrong information communicated
 I3  Information communication incomplete
 I4  Information communication unclear

Selection Errors

 S1  Selection omitted
 S2  Wrong selection made

Planning Errors

 P1  Plan omitted
 P2  Plan incorrect

Violations

 V1  Deliberate actions

Question Set: Identifying human failures

 

Question

Site response

Inspectors view

Improvements needed

1

What does the site understand by the term ‘human failure’?

Do they recognize the difference between intentional and unintentional errors?

     

2

Do they consider that human error is inevitable, or can failures be managed, and how?

     

3

What are the typical ways to prevent human failure?

     

4

What are the main hazards on the site? How has the site addressed human failures that may contribute to major accidents? (e.g. if a significant risk is reactions in batch processes, how has the site addressed human failure in charging incorrect amount or type of product? e.g. if a significant risk is transfer between storage and road/rail tankers, how has the site addressed temporary pipework/hose connection failures?)

     

5

Is there a formal procedure for conducting human failure analyses? – Is there any science/method to how they assess human failures, or is it seen as ‘common sense’?

     

6

Does the site identify those manual operations that impact on major accident hazards? (for example, maintenance, start-up, shut down, valve movements, temporary connections).

     

7

Does the site identify the key steps in these operations? – How (e.g. by talking through the task with operators, walking through the operation, reviewing documentation)? – How do they record this analysis/what formal techniques used (if any)?

     

8

Does the site identify potential failures that may occur in these key steps (e.g. failure to complete the task, completing tasks in the wrong order)?

     

9

What types of failures did the site identify? – Do they include unintentional failures as well as intentional violations? – Do they address mental (decision making) failures or communication failures, as well as physical failures?

     

10

If they claim to perform human failure analyses ‘as part of HAZOP’, what list of potential failures do they refer to (i.e. what is the error taxonomy – does it include action too early, too late, on wrong object, action in wrong direction etc.). – If such a structure is not used then how do they ensure that all potential errors are identified?

     

11

Does the site identify factors that make these failures more or less likely (such as workload, working time arrangements, training & competence, clarity of interfaces/labelling)?

     

12

Has the site considered the hierarchy of control measures in addressing the human failure (e.g. by eliminating the hazard, rather than simply providing training)?

     

13

Do control measures focus solely on training and procedures? – Is there any recognition that people do not always follow procedures? – How do they ensure that people always follow procedures? – What factors do they consider might lead to non-compliance with procedures? – Is there awareness that training can only help to prevent mistakes (mental errors) and that training has no effect in preventing unintentional failures (slips) or intentional violations?

     

14

Do analyses lead to new control measures, or are failures considered to be addressed by existing controls? – Obtain an example of a measure that was implemented as a result of human failure analysis.

     

15

Have attempts been made to optimise the performance influencing factors to make failures less likely (e.g. addressing shift patterns, increasing supervision, updating P&IDs/procedures, clarifying roles)?

     

16

Are operators involved in assessments of activities for which they are responsible? (e.g. task analysis or identifying potential failures).

     

17

How has the site recorded such assessments?

     

18

What training/experience do the assessors have to demonstrate that they are able to identify potential human failures and means of managing them? – How do they know that they have identified all of the failures and influencing factors?

     

19

Have estimates of human failure probabilities been produced? – By what technique? – What were these probabilities used for? – How precise are these estimates and what are the confidence intervals?

     

20

Has the site employed external help/advice in conducting these assessments?

     

21

Has the site considered human failures in process upsets or emergency situations? Have they considered how the influences on behaviour may be different under these circumstances? (e.g. people may experience higher levels of stress in dangerous or unusual situations, or their workload might be greatly increased in an upset).

     

22

Does the analysis focus on operator failure, or do they address management failures? – what about failures in planning, allocation of resources, selection of staff, provision of suitable tools, communications, allocation of roles/responsibilities, provision of training, organisational memory etc.)?

     
Scroll to Top