The problem
Evidence work becomes risky when the system blurs what a source said, what the user corrected and what AI inferred. The product needed to keep those layers distinct while still helping someone organize a difficult body of material.
The working context
The problem
Evidence work becomes risky when the system blurs what a source said, what the user corrected and what AI inferred. The product needed to keep those layers distinct while still helping someone organize a difficult body of material.
My contribution
I am shaping the product, evidence model and approval boundaries. The work focuses on preserving source history, recording user corrections, keeping AI interpretation visible and preventing consequential actions from happening without a person choosing them.
I have kept the organisation, product and people private. I also left out exact figures and internal details that could identify the work. The status above tells you how far the implementation actually went.
How I worked through it
If you are working on a similar system, these are the decisions I would examine before choosing tools or adding more automation.
The source needs to remain available in its original form. Summaries and extracted facts should point back to it so a user can check what changed.
A user correction should not silently overwrite the system's earlier interpretation. Keeping the correction history visible makes later reasoning easier to audit.
The workspace can help organize and draft, but the user remains responsible for actions that could affect their position. Assistance and authority are separate product decisions.
Expert insights from the work
These are the lessons I took from the work. I would use them as questions for your own implementation, not as a universal recipe.
Do not hide provenance in a database field. Show people where a statement came from at the moment they decide whether to trust it.
When your user supplies a correction, record it as a correction. The system should not pretend it learned that fact independently.
A confident answer can still be wrong or incomplete. Decide what the system may do by looking at consequence and user responsibility, not the tone of the generated text.
What needs proving before launch
This product is still in active development and has not launched as a financial service. Before launch, I would test the evidence model with users and relevant subject expertise, especially where a tidy AI summary could hide an important disagreement in the source material.
Connected disciplines
Continue through the work
If this overlaps with a system you are thinking through, start a relevant conversation.