Arc Skills / Safety
Review fault detection isolation and recovery logic
This task reviews a fault-management sequence from occurrence to stable end state. Arc Skills produces a state and timing gap table that identifies where detection or switchover may be too late.
Use this skill
Use Arc Skills to review fault detection, isolation and recovery for the current configuration. Trace representative faults through detection thresholds, transitions, isolation target, alert and stable state.
Inputs:
- Fault list, state machine and controlled timing limits.
- Software/hardware versions and fault-injection records.
- Hazard-effect timing and operator assumptions.
Return a sequence table with expected versus observed transitions, elapsed time and missing boundary cases. Do not rewrite the hazard limit to fit observed recovery.What you provide and what you get
| What you have | How it is used | What you get |
|---|---|---|
| Fault and state definitions | Defines transition path | Sequence map |
| Hazard timing limit | Supplies deadline for effect prevention | Timing comparison |
| Fault-injection records | Shows actual version and boundary behavior | Evidence gap |
A two-second switchover misses a one-second limit
Illustrative engineering example.
A synthetic dual-channel controller has controlled hazard analysis H-3 requiring erroneous output to stop within 1.0 s of sensor timeout. Software V4’s state description shows detection at 1.4 s and switchover complete at 2.0 s. No boundary-case injection result is supplied.
| State or event | Time from fault | Claimed behavior | Finding |
|---|---|---|---|
| Sensor timeout begins | 0.0 s | Fault becomes active | Start of controlled timing interval |
| Fault declared | 1.4 s in V4 design | Failed channel marked suspect | Already 0.4 s beyond 1.0 s limit if output persists |
| Healthy channel active | 2.0 s in V4 design | Output switched | Design sequence exceeds limit by 1.0 s; injection test missing |
The timing comparison is simple: 2.0 s − 1.0 s = 1.0 s. It is a design-level gap, not a measured failure, because no V4 fault-injection run was supplied. Whether the erroneous output actually persists until switchover depends on the isolation logic and must be tested.
The review should challenge both the threshold-crossing boundary and recovery state. A faster fault declaration, independent output inhibit or other approved control might close the gap; selecting one is an engineering decision. The controlled hazard limit remains unchanged until the responsible safety process revises it.
Follow the complete state path and clock
- Define fault occurrence, hazard clock start and safe-state boundary.
- Trace detection, isolation, annunciation, switchover and reset transitions.
- Compare worst relevant elapsed times with controlled limits and phase assumptions.
- Request injected fault tests at thresholds, persistence windows and recovery edges.
Questions about this task
Does a 2-second design transition prove a hazard?
It shows a mismatch with the stated 1-second control basis. The actual output behavior and any independent inhibit still need evidence.
Why test boundary cases?
Threshold and persistence behavior can change which state transition occurs; nominal injection alone may not exercise the limiting path.
Sources and further reading
- NASA Systems Engineering Handbook: Lifecycle, requirements, verification and technical-management guidance.
- NASA Fault Tree Handbook with Aerospace Applications: NASA-hosted guidance for fault-tree gates, cut sets and reliability block diagrams.