Before an AI-assisted mortgage workflow reaches borrowers, the credit union should require seven separate proofs: bounded authority, accurate source handling, fair-lending performance, reproducible reasons, valuation controls, a workable exception path and production monitoring with rollback.
The distinction matters because “mortgage AI” can describe very different tools. One system may classify documents. Another may flag suspected fraud, estimate collateral value, recommend conditions, draft a notice or help an employee review an application. Each use changes a different part of the control environment. A single vendor accuracy figure cannot validate all of them.
The NCUA's current AI guidance says existing regulations remain technology-neutral and that examiners evaluate safety and soundness, compliance, internal controls, ongoing risk monitoring and third-party due diligence. The agency specifically points credit unions to fair lending, data privacy, operational resilience and model risk when considering AI vendors.
Credit unions are already moving AI-assisted decisioning into production in other loan categories. The implementation lesson from Communication FCU's lending deployment is that the operational work begins after model selection: teams still need explicit decision rights, exception ownership and measures that show whether the member outcome improved. Mortgage lending adds collateral, valuation and notice requirements to that foundation.
The following seven tests turn those obligations into a go-live gate. They are a control-design framework, not a substitute for mapping the credit union's specific products and jurisdictions with legal and compliance teams.
1. The decision-rights test
Write down exactly what the system may read, calculate, draft, recommend and execute. Then name the person accountable for each result. Avoid labels such as “AI underwriting” that hide multiple actions inside one phrase.
A useful decision-rights table should identify:
- the workflow step and data the system can access;
- whether the output is a classification, estimate, recommendation, decision or communication;
- the employee role required to review or approve it;
- the evidence that must remain with the loan file; and
- the conditions that force a stop or human handoff.
Start lower on the authority ladder than the technology permits. A document-classification tool may be able to populate fields automatically, but the first production tier might limit it to preparing a review queue. An underwriting model may generate a recommendation, but the credit union can retain the decision with an authorized employee until validation and monitoring evidence justify any broader use.
Pass evidence: a signed workflow map shows one owner, one review rule and one stop condition for every AI-assisted action.
2. The source-to-output accuracy test
Mortgage files contain structured fields, scans, pay statements, tax records, bank statements, appraisals and borrower-supplied explanations. Accuracy should therefore be measured by document type and field, not as one blended percentage.
Build a representative test set that includes clean files and difficult ones: low-quality scans, amended documents, multiple employers, self-employment, nonstandard income, conflicting dates and missing pages. Trace each important output back to its source. When the system extracts or summarizes a fact, an employee should be able to see the document and location that support it.
Measure false acceptance as well as missed automation. A tool that frequently routes good documents to review may slow the process; a tool that confidently accepts the wrong income or identity information creates a more serious control failure. The same principle applies to document automation generally: the useful metric is not extraction speed alone, but whether downstream rework and exceptions decline.
Pass evidence: field-level results meet a preapproved threshold by document and borrower scenario, every material value is traceable, and uncertain or conflicting evidence fails into review.
3. The fair-lending test
Do not treat a vendor statement that protected attributes are excluded as a fair-lending test. Other inputs, interactions and cutoffs can still produce different outcomes. The credit union should compare approvals, denials, pricing, conditions, document requests, overrides and processing time across relevant groups and products.
The NCUA's Equal Credit Opportunity Act overview reminds credit unions that automated underwriting settings must comply with fair-lending requirements. Testing should include the full workflow, not only the final score. A document model that creates more unresolved exceptions for one borrower group can affect access even if it does not make the credit decision.
Define the review before seeing the results: comparison groups, minimum sample, escalation threshold, statistical or file-review method and remediation owner. Small credit unions may need pooled periods, matched-file review or other proportionate methods, but “our volume is too low” is not a monitoring plan.
Pass evidence: compliance can reproduce the test, explain any material difference and document the decision to remediate, constrain or accept the remaining risk.
4. The reasons-and-notices test
An employee must be able to identify the actual principal factors behind an adverse action. The CFPB's circular on complex algorithms states that ECOA and Regulation B apply regardless of the technology and that creditors cannot use model complexity as a reason for failing to provide specific, accurate reasons.
Test this with cases, not a slide deck. For each sample decision, compare the model inputs and factors actually scored with the reason codes and the notice a borrower would receive. Generic labels, a nearest-match checklist or an explanation layer that only approximates the underlying model should not be accepted without validation.
If a tool helps draft requests, conditions or notices, test whether the language is accurate, understandable, accessible and consistent with the loan record. The employee approving the communication should see the supporting facts, not only polished prose.
Pass evidence: sampled notices state the actual principal reasons, trace to the decision record and survive independent compliance review.
5. The collateral and AVM test
If the workflow uses an automated valuation model for a covered mortgage transaction, treat valuation as its own control lane. The interagency AVM quality-control rule, adopted by agencies including the NCUA, calls for policies, procedures and control systems designed to ensure confidence in estimates, protect against data manipulation, avoid conflicts of interest, require random-sample testing and reviews, and comply with nondiscrimination laws.
The implementation question is not whether the vendor reports a low average error. Mortgage teams should define where the model is eligible, what confidence or data conditions trigger another valuation method, how property-type and geography performance are reviewed, and who investigates a borrower challenge or unusual result.
Keep model selection, ordering, review and override evidence. Random samples should reach beyond the easiest properties, and the test set should include thin-data markets and property types that tend to produce exceptions.
Pass evidence: the credit union has a documented AVM control system, independent sample review, escalation rules and a record of how accuracy and nondiscrimination are monitored.
6. The exception-and-human-handoff test
A normal-path demonstration does not show whether the mortgage operation can support the tool. Run scenario tests involving missing consent, conflicting income, an identity or fraud alert, an inaccessible output, a borrower dispute, an unavailable vendor service and a case outside the model's intended population.
The handoff should preserve the work already completed and make the reason for escalation visible. Borrowers should not have to restart because the automation stopped. Employees need authority to pause an output, a queue with a named service level and a route for compliance or technical review.
Measure the exception rate, age, resolution, repeat document requests, borrower contacts, employee overrides and resulting changes. A faster normal path can still be a worse mortgage process if the exception queue grows or difficult borrowers wait longer.
Pass evidence: employees complete the exception scenarios without bypassing controls, losing file context or sending the borrower into an unowned queue.
7. The production-change and rollback test
Validation is tied to a system version, data environment, workflow and set of assumptions. The go-live package should therefore define which changes require notice, retesting or approval: model versions, thresholds, prompts, data sources, subcontractors, integrations and downstream rules.
The NCUA says credit unions using an AI vendor should understand how the product functions, the risks it introduces, how it fits the business model, and the vendor's safeguards, reliability and controls. Translate that expectation into contract and operating requirements. The credit union should receive enough change information to decide whether prior validation still applies.
A production dashboard should combine model and workflow evidence: accuracy, overrides, exception volume, adverse-action reasons, fair-lending indicators, valuation review, complaints, rework, outages and service time. Set trigger levels and name the person who can restrict the system to assist-only mode or roll it back.
Pass evidence: monitoring has owners and thresholds, material vendor changes trigger review, and the team has successfully rehearsed a rollback without losing the loan record.
The minimum mortgage AI evidence packet
Before approving a limited pilot, the mortgage, risk, compliance, technology and vendor-management teams should be able to assemble one evidence packet containing:
- the scoped use case, intended population and decision-rights table;
- the pre-AI baseline and pass thresholds;
- source-to-output validation by document and borrower scenario;
- fair-lending test design, results and remediation decisions;
- sample decision records, reason codes and borrower communications;
- AVM policies and sample review when automated valuations are in scope;
- exception scenarios, queue ownership and service levels;
- vendor due diligence, change-notice terms and incident responsibilities;
- the production dashboard, trigger levels and review cadence; and
- the approval record, rollout boundary and rollback owner.
This packet also creates the training agenda. Employees do not need abstract instruction about every form of AI; they need to know the system's authority, how to verify its evidence, when to override it and how to route a borrower without losing context. The same task-based approach appears in our six-step AI workforce plan.
The final decision rule is simple: a mortgage AI tool should not advance because the average result looks promising. It should advance only when the credit union can show, with reproducible evidence, how it behaves for the normal file, the difficult file, the protected borrower, the adverse-action case, the challenged valuation, the system change and the day it must be turned off.