Ben Squire
Companies House publishes ownership data through a clean, structured API. Directors are there. Persons with significant control are there. Pull the endpoint and you get well-formed, machine-readable fields, ready to drop into any system for seamless processing.
While helpful, two of the most valuable pieces of the ownership picture aren’t there:
- Shareholders. Who owns how many shares, and of what class, sit buried inside confirmation statement PDFs.
- Subsidiaries. The companies a business owns downstream sit buried inside annual accounts PDFs.
Neither has an API field. Neither was ever meant to be read by a machine. The usual result? More manual work for your legal deal team to sort and extract by hand. StructureFlow feels there should be a better way.
Locked in PDFs by design
Confirmation statements and annual accounts are unstructured documents: scanned or generated PDFs, free-form tables, layouts that shift from one company to the next and one accountant to the next. There’s no schema to query. With multiple documents to parse through, all formatted however the filer’s accountant happened to format it, valuable time is inevitably Lost.
That’s a problem for anyone trying to build a complete, structured view of corporate ownership. Directors and PSC data get you part of the way, but without shareholders and subsidiaries, the picture stops short of a full corporate structure.
How StructureFlow’s AI closes the gap
StructureFlow uses Mistral’s OCR and document-understanding model to read these filings the way a human analyst would at a fraction of the time and cost.
The model isn’t asked to summarise a PDF. It’s given a strict JavaScript Object Notation (JSON) schema and a targeted prompt for each document type, and it returns validated, structured data: shareholder name, number of shares, share class; subsidiary name, country, registration number.
What goes in as a scanned legal filing comes out as machine-usable data, ready to slot straight into a company’s ownership graph on your behalf.
Why this needed AI, not just automation
Reading every filing by hand doesn’t scale. It’s slow, it’s expensive, and it can’t be done reliably across thousands of companies.
Traditional parsing – regex, fixed templates – doesn’t hold up either. The moment a filing’s layout changes, the extraction breaks, and every accountant formats these documents slightly differently. Rule-based extraction is built for consistency that doesn’t exist in the real world of company filings. A document-understanding model generalises across formats in a way fixed rules never can.
The result: structured API data (directors, PSC) combines with AI-extracted data (shareholders, subsidiaries) to produce one full corporate structure, in one API call. Finally, you have an easily referenceable, hierarchically structured company diagram with shareholders and subsidiaries.
Built for accuracy, speed, and cost
Two design choices keep the extraction sharp and the economics sensible.
- Most-recent-first. Filings are read newest to oldest, so the shareholders and subsidiaries extracted reflect where the company stands today, not where it stood several filings ago.
- Smart stopping. The service reads filings only until it finds the first document that yields usable data, then stops. Cost and latency stay bounded, without sacrificing the currency of the data.
Together, these turn a stack of inconsistent legal PDFs into something a platform can actually use: a complete, current, structured ownership picture built in seconds, not hours.
Explore StructureFlow
Ready to unlock speed and automation at your own firm? Book your 20 minute demo of StructureFlow today.




