Financial Services
Document processing automation with generative AI
A financial services provider automates onboarding document handling with grounded extraction, human review and a complete audit trail.
- Sector
- Financial services
- Scale
- High-volume client onboarding and periodic review
- Engagement
- AI readiness workshop and phased implementation
- Duration
- Approximately eight months
The challenge
Where this started
Onboarding required operations staff to read incorporated documents, statements and identification, key fields into core systems and file evidence. The work was slow, repetitive and inconsistent between reviewers, and volume peaks translated directly into backlog.
Automation had been considered and shelved. Compliance would not accept a system that produced answers without showing where they came from, and nobody could explain what data would be sent to which model or how a regulator would audit a decision months later.
Our approach
What we did, and why
The decisions that mattered, including the ones that were unglamorous.
- 01
Settle the data boundary first
Before any build, we documented which document classes could be processed, where they could be processed, what would be retained and for how long — and had compliance agree it in writing. That unblocked a project that had previously stalled at review.
- 02
Extraction grounded in the source
Every extracted field carries a citation back to the document and page it came from. Reviewers see the evidence beside the value, which is what made the output acceptable to compliance.
- 03
Confidence routing with humans in the loop
High-confidence extractions on straightforward documents flow through with spot checks. Anything ambiguous, low-confidence or high-value routes to a specialist. The system reduces reviewer load rather than removing the reviewer.
- 04
Evaluate before every release
A labelled set of representative documents, including deliberately awkward ones, scores field-level accuracy automatically. A release that regresses on any critical field does not ship.
- 05
Audit as a first-class output
Every extraction, confidence score, routing decision and human override is logged immutably, so a case can be reconstructed exactly as it was decided.
Architecture
How the pieces fit together
01
Intake
- Document upload
- Classification
- Cloud Storage with retention policy
02
Preparation
- Text and layout extraction
- Chunking
- Metadata and permissions
03
Retrieval
- Vector search
- Permission-aware filtering
- Re-ranking
04
Model layer
- Vertex AI
- Gemini
- Prompt and context management
05
Review
- Confidence routing
- Specialist review queue
- Immutable audit log
Technologies
- Vertex AI
- Gemini
- Cloud Storage
- BigQuery
- Dataflow
- Cloud IAM
- Cloud Logging
Outcomes
What changed
Described qualitatively. We publish client-specific figures only where the client has approved them.
Manual keying substantially reduced
Straightforward documents flow through with spot checks, concentrating specialist attention on genuinely difficult cases.
Peaks absorbed without backlog
Volume spikes that previously created queues are handled by throughput that scales with demand.
Consistent treatment
The same document produces the same extraction regardless of who is on shift, removing reviewer-to-reviewer variance.
Defensible under audit
Citations and an immutable log mean a decision can be reconstructed and explained months later.
Compliance approved the design
Agreeing the data boundary before building turned the review from an obstacle into a checkpoint the project passed.
This engagement is anonymized and presented as a representative scenario. It reflects the architecture patterns, decisions and trade-offs typical of our work in this area rather than the details of one named client. Identifiable engagement details are published only with written client approval.
Related services
Other engagements
- Enterprise data warehouse migration to BigQuery
Retail
- Predictive maintenance on Vertex AI
Manufacturing & Energy
