Introduce AI into systems engineering through one bounded workflow with a known owner, controlled data and a reviewable result. For a space team, the first pilot should establish whether assistance improves a specific engineering task under real review conditions. A short adoption plan needs explicit success criteria, difficult test cases and a decision to continue, narrow, redesign or stop the work.
Start with a decision the team already makes
Choose a recurring question such as which verification records need attention after a payload requirement changes. Identify the current inputs, people involved, expected output and handover. Describe where work is difficult: locating the applicable interface, reconstructing why a trace link exists, or checking whether an earlier test result still applies. That becomes the pilot's problem statement.
Resist choosing the most expansive task available. A pilot for preparing an impact note has a clearer review boundary than one for designing an entire subsystem. Keep the ordinary engineering process visible throughout. NASA's requirements-management guidance connects changes with impact assessment and approval; the pilot tests how assistance contributes to that work.
Choose the simplest mechanism that can answer the question. A rule may detect a missing mandatory field, a query may collect existing links and a language model may help interpret ambiguous context. Anthropic's implementation guidance favours starting simply and evaluating added complexity. Its principles are useful here without assuming that a space team needs any particular agent framework.
Check readiness before processing programme records
Confirm the data owner, permitted processing environment and applicable access rules. A record being available to one engineer does not automatically make it suitable for every model route. Establish where source data, generated outputs and logs will reside, who may retrieve them and what retention arrangements apply. Resolve those questions for the pilot's actual configuration.
Next, check engineering readiness. Can the team distinguish the approved baseline from proposed revisions? Do requirements, interfaces and verification records have stable identifiers? Can a reviewer open the relevant evidence? If those conditions fail, spend the first stage repairing the chosen record set. The pilot should not quietly become an organisation-wide data-cleaning project.
Finally, confirm that the team can review the output. Name the workflow owner, technical reviewers and person authorised to suspend the pilot. Reserve time for inspecting both useful and incorrect suggestions. NIST's AI Risk Management Framework organises work around governance, context, measurement and risk management. These readiness questions apply that structure to a small engineering evaluation.
A reusable pilot charter
Adaptable planning template. Complete this charter before observing the evaluation results. The example entries describe a proposed pilot, not an executed Arc deployment. Replace them with the team's actual scope and record the reasoning behind its criteria.
| Charter field | Example entry |
|---|---|
| Engineering question | Prepare a payload-power change impact note for review against one baseline. |
| Permitted input | Approved snapshot of requirements, related interfaces, verification plans and relevant analysis assumptions. |
| Permitted actions | Read the snapshot, run approved calculations and prepare a separate draft with source references. |
| Deliverable | Confirmed conflicts, affected records, numerical checks, missing inputs and a proposed review route. |
| Decision owners | Assigned systems, payload, electrical power and verification roles under the existing programme process. |
| Evaluation comparison | Current process and assisted process on equivalent cases, including review and correction effort. |
| Stop conditions | Unauthorised access or changes, fabricated consequential evidence, or inability to reconstruct a decision. |
| End decision | Continue within the same scope, narrow the task, repeat after correction or stop. |
The charter should also identify the assistant configuration, retrieval settings and relevant tool versions. If the team changes them during the pilot, preserve the earlier results and distinguish the new evaluation. Otherwise, a final average may combine different systems and conceal whether the remedy actually improved the failure it addressed.
An illustrative 30-day plan
The calendar below provides an organising window for a team that already has access to a suitable environment and reviewers. It is not a promised implementation duration. If data preparation or a consequential defect takes longer, revise the schedule and the decision date rather than skipping the evaluation.
| Period | Work | Exit evidence |
|---|---|---|
| Days 1 to 5 | Select the question, assign owners, confirm data boundaries and document the current workflow. | Completed charter, accessible record snapshot and agreed evaluation criteria. |
| Days 6 to 10 | Prepare reference cases and configure a bounded draft workflow. | Reviewed expected findings, recorded assumptions and demonstrated permission boundaries. |
| Days 11 to 20 | Run cases, inspect errors and compare complete review tasks. | Case-level findings, corrections, time records and separately recorded failure categories. |
| Days 21 to 25 | Exercise difficult conditions, ordinary handovers and fallback. | Results for missing sources, changed revisions, ambiguous context and recovery. |
| Days 26 to 30 | Review the evidence and make the scoped adoption decision. | Signed-off internal decision record, remaining limitations and next owner. |
During the middle stage, let engineers complete ordinary tasks with the assistant in a draft or shadow arrangement appropriate to the programme. Record intervention rather than hiding it. If the assistant succeeds only because a reviewer supplies an omitted interface, that is a useful observation about the workflow, not an unassisted success.
Worked pilot data: build cases around credible mistakes
Fictional evaluation example. Use a 6U Earth-observation CubeSat with PAY-PWR-014 revision B in baseline BL-03. The requirement caps payload peak power at 20 W. The design currently uses 18 W and a proposal raises it to 24 W. Include interface ICD-EPS-PAY-02 and verification activity VER-PWR-07. The expected findings below are authored reference answers, not results observed from a model.
For the base case, a correct impact note distinguishes the 6 W design increase from the 4 W cap exceedance. If the programme has a positive power margin of 5 W under the same operating condition, keeping everything else fixed produces negative 1 W. It identifies interface and verification reassessment and does not declare the proposed design acceptable. The technical reviewers should approve this reference interpretation before using it to score outputs.
Create separate case variants without altering the base case's meaning. In one, remove access to the interface and expect an explicit limitation. In another, supply an obsolete copy of the requirement alongside the applicable revision and check revision selection. In a third, label a value as orbit-average power and check that it is not substituted for the peak constraint. Include a clean case in which no conflict should be invented.
Add a variant whose answer is genuinely uncertain, such as an undefined measurement boundary. The reference outcome should request clarification and withhold the dependent conclusion. This prevents an evaluation from rewarding confident completion when abstention is the useful engineering behaviour. Also test whether a misleading instruction embedded in a source document can cause the assistant to disregard its task or permissions.
These case types form a coverage checklist, not a statistical sample-size prescription. Select enough independently reviewed examples to exercise the decisions and failures that matter in the chosen scope. Rephrasing one case many times does not establish broad coverage. A small pilot can expose defects and support a local decision; it cannot prove a rare-error rate across an entire mission lifecycle.
Measure the complete engineering task
Keep a case-level evaluation record with the input snapshot, expected finding, actual output, reviewer disposition and correction required. Separate a missed conflict from an unnecessary warning, an invented citation and a wrong revision. A single quality score can obscure those differences and encourage teams to trade a consequential miss for many easy correct answers.
Measure source accuracy by opening the cited record and checking both its applicability and support for the claim. Measure useful abstention on cases with missing information. Record whether affected records were found and whether the final package distinguishes proposed work from approved evidence. Have a technical reviewer settle disputed reference answers before counting the assistant as wrong or right.
For effort, include data preparation, task setup, execution, engineer review, correction and handover. Compare equivalent cases and describe how prior familiarity could affect timing. An engineer who has already solved a case may review it unusually quickly. If the comparison cannot control that effect, report it as a limitation instead of presenting the difference as a general productivity gain.
NIST's Generative AI Profile, particularly its measurement guidance, cautions against extrapolating from narrow assessments. Report what the pilot actually tested, including the human support it needed. Keep the fictional power figures separate from observed evaluation results.
Set the go or no-go criteria before scoring
Some criteria are categorical. The pilot must remain inside its access and change boundaries, and consequential claims must be traceable to inspectable support. An observed boundary violation warrants stopping the affected workflow and investigating it. Passing these checks in a small trial does not prove that no future violation is possible.
Other criteria need programme judgement. Agree in advance which issue types must be detected in the scoped cases, what reviewer burden is acceptable and which uncertainty must trigger escalation. Base those choices on the decision's consequence and the current process. There is no universal percentage of accepted suggestions that makes an engineering assistant suitable.
At the final review, choose a disposition with a specific scope. Continue if the evidence supports the chartered task; narrow the workflow if only part is useful; repeat relevant evaluation after a fix; or stop if the task cannot be reviewed reliably. In the fictional CubeSat pilot, correctly identifying the cap conflict is necessary, but inventing a passing result for VER-PWR-07 would still prevent a favourable decision.
If use continues, assign ownership for monitoring, configuration changes and fallback. Reassess when the model route, retrieval behaviour, tools or programme data change materially. Retain a compact set of difficult reference cases so the team can check whether a new configuration reintroduces a known failure. Expansion should answer a new scoped question with fresh evidence.
Using Arc as the programme context for a pilot
Arc keeps requirements, architecture, systems, interfaces, risks, verification activities and evidence in a connected programme model. Its read-only and drafting agents can surface context, check records and propose work within configured permissions. Those capabilities can support a bounded evaluation around accessible programme records.
Arc also supports branches, record differences, discussion, review and merge before changes enter a controlled baseline. The pilot still needs the team's authority model, appropriate deployment configuration and reviewed evidence. Its deliverable should be a defensible adoption decision about the tested workflow, with limitations another engineer can inspect.
Frequently asked questions
How long should a systems engineering AI pilot last?
Long enough to prepare representative data, evaluate the workflow and reach an evidence-based decision. The 30-day schedule in this guide is an illustrative planning window. A pilot should extend or stop if unresolved data, review or technical issues prevent a sound decision.
How many test examples does an AI engineering pilot need?
There is no universal minimum. Choose cases around the intended decisions and credible failures, then judge whether important conditions have been covered. A small set of worked examples can debug a workflow but cannot establish a rare-error rate or programme-wide reliability.
Should the first pilot use real spacecraft programme data?
Use data the team is authorised to process in the chosen environment. Fictional or suitably prepared historical examples can establish basic behaviour. Before making operational claims, evaluate representative authorised programme material with its actual revisions, permissions and review conditions.
What should be measured beyond response speed?
Measure correct findings, missed issues, false alarms, source and revision accuracy, appropriate abstention, reviewer correction effort and whether the final disposition is supported. Include preparation, review and recovery effort when comparing workflows.
When should a team stop or redesign an AI pilot?
Stop the affected workflow if it violates access or change boundaries, fabricates consequential evidence or cannot be reviewed reliably. Redesign or narrow the task when repeatable errors exceed the programme-defined criteria, and retest the relevant cases before expanding use.
Evaluate Arc
Try Arc on a representative engineering workflow
Start with one requirement set and test traceability, change control, review and verification in a private Arc workspace.