Credit unions should review every AI initiative on five evidence lanes—member and mission outcome, realized economics, control performance, operating readiness, and dependency and exit. Scale only when all five support broader use and no hard-stop condition is present. Put uncertain initiatives on a time-boxed hold. Stop initiatives that cannot meet a mandatory control, demonstrate value or exit safely.

This is a portfolio decision, not another business-case template. A business case asks whether one proposed use is worth funding. A portfolio review asks where the credit union should place its next dollar, its limited risk capacity and its scarce implementation attention across pilots, production systems and experiments already under way.

The need for that discipline is growing. The NCUA’s AI resource, updated April 28, 2026, says credit unions should understand how a product functions, the risks it introduces, how it fits the business model, and the vendor’s safeguards, reliability and controls. NCUA also says examiners use the existing supervisory framework, including internal controls, ongoing monitoring and third-party due diligence. The Treasury Department’s February 19, 2026 financial-services AI framework release likewise emphasizes evaluating use cases and managing risk across the lifecycle.

Neither source tells a board which local initiative deserves the next tranche of funding. The scorecard below turns those lifecycle expectations into a repeatable allocation decision.

Start with three decisions, not one blended score

A single weighted total can hide the exact fact a board needs to see. Strong projected savings should not cancel a failed fair-lending test, and good user adoption should not offset an inability to recover records from a vendor. Keep each evidence lane visible and use three decision states:

  • Scale: the initiative met its approved outcome and control thresholds at representative volume; the operating owner can support the next stage; and funding, monitoring and fallback capacity are available.
  • Hold: value remains plausible, but one or more remediable evidence gaps block expansion. The hold names the gap, owner, budget and deadline—normally 30, 60 or 90 days.
  • Stop: a hard-stop condition is triggered, the time-boxed hold expires without sufficient evidence, or the initiative no longer outranks better uses of money and attention.

A hold is not permission to drift. If leaders cannot specify what evidence would change the decision, the initiative belongs in stop or scale.

Lane 1: Member and mission outcome

Ask whether the system changed the outcome it was funded to improve. Measures should follow the complete member or employee task: resolution time, repeat contacts, avoidable declines, loss prevented, accessibility success, error correction, employee capacity redirected to a named service, or another mission-linked result.

Scale signal: a measured outcome improves for the intended population without a worse result being shifted to another channel or member group. Hold signal: output quality looks promising but the outcome window or comparison group is incomplete. Stop signal: the initiative produces activity without an attributable outcome, degrades member experience or depends on a benefit the operating owner cannot realize.

Lane 2: Realized economics

Replace the original forecast with actual cost and realized value. Include integration, data work, validation, monitoring, employee review, exceptions, remediation, training, vendor changes and exit—not only license fees. Separate cash impact, released capacity and member or control outcomes so one improvement is not counted three times.

Our AI business-case framework explains how to baseline one use case and fund evidence in stages. At the portfolio review, compare actual results from that benefits ledger. Scale when the base case survives with observed inputs. Hold when a specific adoption or integration barrier has a credible fix. Stop when the fully loaded cost persistently exceeds the approved value range or the claimed capacity has no redeployment plan.

Lane 3: Control performance

Bring the evidence packet, not a statement that the model “performed well.” It should show the version in use, scenarios tested, error and override rates, complaints, security events, privacy or fair-lending tests where applicable, human escalation, corrective actions and unresolved exceptions.

The NIST AI Risk Management Framework organizes work around govern, map, measure and manage. Its companion AI RMF Playbook treats risk management as iterative rather than a one-time approval. That supports a simple portfolio rule: scaling is a new risk decision. A pilot’s controls must be re-tested against the broader population, data, volume, access and consequence of the next stage.

Scale only when control thresholds are met and owners can monitor them at the next volume. Hold a fixable validation, documentation or monitoring gap. Stop when a mandatory legal, safety, security or member-protection control cannot be met.

Lane 4: Operating readiness

Measure whether the workflow works around the tool. Look at eligible volume, employee adoption, override and correction work, exception queues, supervisor capacity, member handoffs, incident coverage and service continuity. High model accuracy can coexist with a poor operating design.

Scale signal: the next-stage owner has staffing, procedures, training, fallback and monitoring capacity. Hold signal: one operational bottleneck is understood and funded. Stop signal: expansion would create unmanaged exceptions, remove a necessary human review or make critical service dependent on an unsupported process.

Lane 5: Dependency and exit

A use case can create value and still be the wrong portfolio bet if it locks the credit union into an opaque vendor, duplicates another capability or cannot be unwound. Record data portability, retained evidence, model and prompt dependencies, integration concentration, replacement options, minimum service levels and shutdown cost.

Use the tests in our AI vendor exit playbook before expanding scope. Scale when the institution can preserve records, continue essential service and exercise contractual exit rights. Hold while a specific portability or continuity gap is corrected. Stop when the vendor cannot support required oversight, the capability is redundant, or broader deployment would make safe exit materially harder.

Define hard stops before the review

The board or delegated committee should approve hard-stop conditions in advance. Tailor them to the use case and applicable law, but common conditions include:

  • a material member-harm, compliance, security or privacy threshold is breached;
  • management cannot reconstruct a consequential output or preserve required evidence;
  • a required human fallback or appeal route is unavailable;
  • the system operates outside its approved data, authority or population;
  • a vendor change prevents required testing, monitoring or audit access; or
  • the approved hold expires without the named evidence.

The NCUA’s September 30, 2025 internal AI compliance plan is not a rule for credit unions, but its termination sequence is a useful public example. NCUA says a noncompliant internal AI system would be restricted or isolated, associated data and assets archived, the rationale documented and relevant stakeholders notified. A credit union can adapt that sequence to its own obligations and vendor contracts.

Make the stop decision operational

“Stop” should launch a controlled retirement plan, not simply end the subscription. Name the shutdown owner and sequence: contain new use, activate the approved fallback, preserve decision and monitoring records, notify affected internal teams and vendors, remove access and integrations, return or delete data as required, confirm member remediation if necessary, and close the contract and post-implementation review.

For a system that is safe but no longer valuable, the sequence can be scheduled. For a material control failure, containment and fallback may need to precede the committee meeting. The portfolio record should capture the trigger, authority, date, residual obligations and lessons for future procurement.

The quarterly portfolio packet

A useful board or executive packet can remain compact. Include:

  • one inventory row for every pilot and production AI use, including owner, vendor, stage, spend and next decision date;
  • the five evidence lanes with current result, target, source, trend and accountable owner;
  • any hard-stop condition, open exception or overdue evidence;
  • the requested scale, hold or stop decision and the resources it releases or consumes;
  • a portfolio view of duplicate capabilities, vendor concentration and total monitoring capacity; and
  • the decision log, including conditions attached to the next funding release.

The board does not need to choose model thresholds or inspect every test case. It should be able to see whether management has defined the outcome, verified the economics, met the controls, prepared the operation and preserved a safe exit—and whether the portfolio still reflects the credit union’s strategy.

A disciplined stop decision is not a failed innovation program. It is evidence that the program can distinguish learning from value and experimentation from permanent overhead. The credit union that can retire a weak AI use case cleanly is better prepared to scale the strong one.

Turn AI pilots into portfolio decisions. Subscribe to the CreditUnionAI Weekly Briefing for practical implementation and governance coverage.

Get the Weekly Briefing