Shutdown overruns and the common culprits
Months of careful planning go into preparation for shutdown events – but overrun is not uncommon in the space. The question arises: how did we end up at this result? It’s rarely a single driver, and more often a reflection of the many moving parts and their incremental impact. In this article we will explore some of the common culprits and their contributions towards a late finish.
Critical and near-critical sequences
A project’s baseline critical path is the collection of activities and their logical requirements which form a driving sequence to set the overall timeline of the shutdown. These sequences receive additional scrutiny and setup to ensure a favourable run and are monitored closely in execution. Their durations are set with a reference point of 0 float (or slack) and all other sequences are measured against it.
Near-critical paths are the next handful of competing sequences closest to a critical path as calculated by their float.
When the critical path cannot be maintained to plan, the cause is rarely the critical path itself; it is usually a symptom of greater issues. Items such as planned productivity, missed interactions, emergent work and more are expanded on in the sections below.
Unidentified coordination
Unidentified coordination has three common sources: (1) interactions between maintenance tasks, (2) from an operational perspective between the systems and (3) technical windows such as commissioning. The most common is maintenance coordination due to volume.
The rate of planned maintenance typically falls into a bell-curve shape with peak rates observed around the midpoint of the event as the workfronts become available. Within the range of potential outcomes, some workfronts may move ahead of their planned projections while others observe slippage. This can result in overlap of interactions that were not previously considered.
Operationally, missing key interactions between the structural phases of systems when shutting down or starting up systems will provide illegal solutions with incorrect float. Systems interactions should be mapped to completeness to eliminate the potential of one system’s drift having an unexpected impact on the project timeline.
Misalignment of de-commissioning and commissioning activities occurs when the work has not been logically entwined into the correct commissioning window. Does this work occur while the equipment is isolated or de-isolated? At startup or under system load? If these are not mapped properly it can result in unexpected interactions and delays where the workflows get held up by an incorrect sequence.
Productivity & resource loading
Productivity factor is the term for the rate of planned work per resource per shift and should be selected with consideration of:
- a) the maintenance planning procedures – has the work been consistently scoped to ‘full shift capacity’, ‘tool-time only’ or somewhere in-between?
- b) the layout and environment of the site – how far are the cribs from the workfronts? Are the workfronts generally at grade or reached by stairs at elevated platforms? How accessible are hydration stations and amenities from the workfronts? What range of temperatures are expected during execution?
If the work is planned at an unachievable rate across the board, key project targets shift from being driven by critical paths to resource limit constraints. When this occurs, the number of sequences that become tied in duration with the driving path becomes unmaintainable. Subsequently, reaching mechanical completion and returning all systems at 0 float can drive an unachievable amount of parallel work for the plant operators to return the plant to service. Both examples identified are expected to drive slippage.
Modelling the entire shift as productive hours would yield a tighter planned timeline but is likely to diverge from realistic expectations unless the work has been planned to reflect this.
Critical resources and technical resource crews can carry an implicit second stage constraint beyond the resource limit itself. Examples of this could be technical crews that require the presence of a leading hand or a site with hazardous operations prescribing tighter requirements for the training and allocation of permit holders. These factors drive an underlying limit on the number of concurrent activities that can be performed simultaneously for the trade.
Schedule adherence
The schedule is frozen, parts are kitted and the mobilisation plan is complete – but what happens when your workforce falls short of what was planned? Sites typically engage contracting partners to bolster their numbers for events; however, a transient, contracted workforce is not a guaranteed workforce. While most individuals make good on their committed rotations, their situation can change and a fraction simply will not arrive for personal or other reasons.
It is wise to identify your critical and potentially constraining resources and plan for additional coverage above the demand with the expectation that some will fall off. Similarly for key equipment such as slew cranes and cleaning trucks, ensure suitable access to replacement vehicles or parts is made available to address typical sources of failure. Where these risks are identified, discuss back-up coverage plans with the relevant contracting partners.
During execution, adherence is a different matter. The clearer it can be made to the supervision around what the group’s priorities are as well as their individual priorities, the greater the chance of maintaining planned timelines and productivity. This is best communicated to supervisors with interaction meetings and runsheets to ensure priority, proximity and expected timings are well known to all.
Addition to scope
Emergent work is the addition of scope after its freeze and can be identified in the lead-in to execution or during execution itself. When identified, reliability and risk assessment should be conducted to ensure the works are necessary, the parts are available and there is sufficient resource capacity to absorb this labour. Scope creep is a known driver of overrun and the works should be limited to those that must proceed in the event.
Where a job’s scope cannot be fully determined prior to the event, such as internal inspection & repair, planners should consult past inspections and historical requirements of the vessel or similar to determine a suitable and conservative allowance to complete these works based on best understanding of condition.
Risk & contingency
Inclement weather conditions during a shutdown will pause many workfronts where the change of environment increases the risk of the task. Cranes typically cease operation when lightning is observed and exposed workfronts are expected to pause during rainfall due to worker discomfort, reduced visibility and greater chance of slip and trip hazards. The event’s location, calendar month(s) and planned duration should be assessed against historical weather patterns for the region to determine suitable weather contingencies.
The larger the scale of the project, the more consideration should be given to risk modelling of a solved schedule to determine probabilistic outcomes of execution. This is completed by assessing individual task-based risks of critical and near-critical sequences and processing Monte Carlo simulations to log the ranges and probability of potential completion.
Incidents
As in all forms of construction and maintenance, safety incidents are never acceptable in shutdowns. A key focus is to ensure all personnel leave site each day with the same health that they came in with. Incidents and any identified unsafe conditions result in investigations and exclusion zones which interrupt execution until such time it is determined safe for the workforce to resume.
Quality incidents can cause rework as a best case scenario and introduce risk to the asset as the worst case scenario. Layers of protection exist to identify these issues such as work packs, QA personnel, permits and PSSR. When identified it carries a cost: an increased demand on resources to redo the works, additional consumables, assurance rechecks and potential for investigation. The later a quality issue is identified, the greater the impact it carries. Poor quality identified on start-up results in winding back of system start-ups to re-establish isolation for intrusive works.