Know the estate. Control every move. Prove every file.
Mergiva discovers and classifies an acquired estate, gates every move on the deal’s stage and two signatures, and re-hashes every file at the destination.
See what the estate holds before anything moves.
Discovery tells you what exists. Classification tells you what it is. Together they turn an unknown estate into a list of files you can scope, sign for and move in batches.
Discover
- Scans Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online, SFTP and local or NAS sources.
- Records each file’s SHA-256, size, content type and last-modified date.
- Feeds the Data Estate Report, a PDF built from the scan’s real data.
- A scan that fails shows its reason, so your team can fix it and run it again.
- Allowed from pre-deal through the TSA period, so diligence can start before signing.
Classify
- AI-assisted by default. Microsoft Presidio redacts detected personal data before a sample leaves your installation.
- You choose the model for each AI feature, and whether calls go to Anthropic with your key or through your own AI gateway.
- A wave’s file scope can be narrowed by classification category and sensitivity, as well as by path, file type and size.
Data Estate Scan
Source: SFTP| File | SHA-256 | Category | Sens. |
|---|---|---|---|
| CSR_0142_final.pdf | 9f3a…c1e7 | GxP-Clinical | 4 |
| patient_listing_v3.xlsx | 2be0…7d19 | PHI | 5 |
| batch_record_BR-2231.pdf | 4b08…d2aa | GxP-Manufacturing | 4 |
| HSR_filing_notes.docx | 61c4…09fe | Legal-Antitrust | 5 |
| FY25_forecast.xlsx | c3d7…5a20 | Financial | 3 |
| site_photos.zip | 8e19…b3c4 | Other | 1 |
AI classification, under your control.
Mergiva uses a large language model to place every discovered file in your taxonomy and score its sensitivity. You choose the model for each AI feature, the account it runs under and how personal data is handled.
In a life-sciences deal the first questions are not about bytes. Where are the clinical records? Which files hold personal or health data? What is competitively sensitive, and must stay behind the antitrust gate until Day 1? The answers decide how the estate is split into waves, who signs for each one, and what your privacy and quality teams need to see.
Mergiva answers them file by file. Once discovery has recorded and hashed every file, a large language model reads each file’s details, places it in your taxonomy and scores its sensitivity from 1 to 5. Every answer is stored with the model’s confidence and its reasoning, so a reviewer can see why a file was tagged the way it was.
You stay in charge of the AI. You choose the model and the account it runs under, how personal data is handled before a request leaves your cluster, and how much file content the model may see. If your policy is that no personal data leaves, set the guard to refuse any classification request that contains it.
How one file is classified
Six steps, and five of them run inside your cluster.
Step 1
Discovered file
Name, path and type
Step 2
Prompt from your taxonomy
Your categories and scale
Step 3
Personal data guard
Redact, block or warn
Step 4
Your chosen model
Claude, or your gateway
Step 5
Checked answer
Only your tags accepted
Step 6
Stored with the file
Tags, score, confidence
Step 1
Discovered file
Name, path and type
Step 2
Prompt from your taxonomy
Your categories and scale
Step 3
Personal data guard
Redact, block or warn
Step 4The one outside call
Your chosen model
Claude, or your gateway
Step 5
Checked answer
Only your tags accepted
Step 6
Stored with the file
Tags, score, confidence
- 1
Discovered file
The discovery scan has already recorded each file’s name, folder path, content type and SHA-256. For files on a local or NAS source, the classifier also reads up to the first 4 KB.
- 2
Prompt from your taxonomy
The classifier builds the prompt from your taxonomy profile: every category with its definition, and the sensitivity scale. It asks for one answer in a fixed format.
- 3
Personal data guard
Before the request leaves your cluster, Microsoft Presidio looks for personal data. By default it redacts what it finds. You can make it refuse the call instead, or only log the finding.
- 4
Your chosen model
The request goes to the model you chose for classification, Claude Haiku 4.5 unless you set another. It travels to Anthropic under your own key, or through your own AI gateway.
- 5
Checked answer
The answer is parsed and checked. Tags outside your taxonomy are dropped, and the score must sit on your scale. If the answer fails the check or the call fails, the file is marked as a fallback with zero confidence, so no one mistakes it for the model’s view.
- 6
Stored with the file
The result is stored against the file: its category tags, a sensitivity score, the model’s confidence and reasoning, and whether it came from the model or the fallback.
Exactly what goes into the prompt
What the model receives depends on where the file lives. The prompt holds nothing else: no deal name, no hashes, no other files.
- Local folders and NAS shares
- File name, folder path, content type, and up to the first 4 KB of the file. Text files are sent as text; other files as a short hexadecimal excerpt.
- Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP
- File name, folder path and content type. No file content.
- Every request
- Your taxonomy’s categories with their definitions, the 1 to 5 sensitivity scale, and the format the answer must take. By default, detected personal data is redacted from the file’s details before the call.
Classification requestLeaves your cluster
- file name
- patient_listing_v3.xlsx
- path
- /clinical/site-07/
- type
- spreadsheet
- sample
- Patient [REDACTED], visit 3 …
- taxonomy
- 8 categories, scale 1 to 5
The sample is sent for local and NAS files only, up to 4 KB.
Answer, stored with the file
- sensitivity
- 5 · Highly Restricted
- confidence
- 0.94
- reasoning
- patient-level clinical listing
Your model, your account, your rules
Each setting below exists in the product today. The note under each says where it is set.
- Your model, feature by feature
- Choose the model for each AI feature, classification included. Claude Haiku 4.5 is the default, and Claude Sonnet 4.6 and Claude Opus 4.6 are also offered. Behind your own gateway, you list the models it serves.
- AI Operations screen, by an administrator
- Your own account, or your own gateway
- Calls run under the API key you give the installation, so they fall under your own agreement with Anthropic. Or point Mergiva at your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. Every call goes only where you route it. A call your settings cannot serve is refused, never sent somewhere else.
- Installation settings
- How personal data is handled
- Redact detected personal data before the call (the default), refuse any request that contains it, or only log it. A second, optional check scans the model’s answer and redacts any personal data it repeats.
- Installation settings
- Your taxonomy
- The categories, their definitions and the names of the five sensitivity levels come from a taxonomy profile. The default is the pharma profile shown on this page. You can supply your own; the scale stays 1 to 5.
- A profile file, read when the service starts
- How much content the model sees
- Set how much of a local or NAS file goes into the prompt, from the default of 4 KB down to none.
- Installation settings
- Rate limits and spend budgets
- Cap requests and tokens per minute, concurrent calls and the size of a single request. Set daily or monthly spend budgets for each tenant. Both are checked before every classification call.
- By a platform or tenant administrator
- A circuit breaker
- When the model keeps failing, Mergiva stops calling it until it recovers, so an outage fails fast instead of stalling every scan. Calls refused for rate limiting are retried with a growing delay.
- Built in
- Usage you can see
- The AI Operations screen shows calls, tokens, cost, cache hits and latency for classification, with the model each call used.
- AI Operations screen
Eight categories and a five-level scale
The categories and definitions the model is given, exactly as the default pharma profile states them. Replace them with your own profile if your organisation classifies differently.
- GxP-Clinical
- clinical trial data, protocols, case report forms
- GxP-Manufacturing
- batch records, deviations, equipment validation
- GxP-Regulatory
- FDA/EMA submissions, 510(k), CMC, CTA, labeling
- PHI
- personal health information (named patients, MRNs, diagnoses)
- PII
- personal information without health context (HR records, contractor names)
- Financial
- revenue, accounting, internal financials
- Legal-Antitrust
- HSR, FTC/DOJ filings, antitrust counsel work product
- Other
- none of the above
Sensitivity, scored 1 to 5
- 1Public
- 2Internal
- 3Confidential
- 4Restricted
- 5Highly Restricted
A file can carry more than one category, such as clinical and health data together. It always carries exactly one sensitivity score.
Classification shapes the rest of the deal
Scoping waves
A wave’s file scope can be narrowed by category and sensitivity, so clinical records, personal data or financial files can move in their own approved batch.
The Data Estate Report
A Data Estate Scan runs discovery, then classification, then builds the PDF report, so the deal team sees what the estate holds before anyone plans a move.
The Classification Explorer
Reviewers filter and browse the results, and open any file to see its tags, score, confidence and the model’s reasoning.
What buyers ask about the AI
Which AI model does Mergiva use?
Claude Haiku 4.5 by default. Claude Sonnet 4.6 and Claude Opus 4.6 are also supported, and an administrator can choose the model for each AI feature. Behind your own gateway, you list the models it serves.
Can we use our own AI account?
Yes. The installation calls the model with the API key you give it, under your own agreement with Anthropic. Or route every call through your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. A call your settings cannot serve is refused, never sent somewhere else.
What exactly does the model see?
For files on local or NAS sources: the name, folder path, content type and up to the first 4 KB. For Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP: the name, path and type only. By default, detected personal data is redacted before the call.
How can we check what the model decided?
Each result carries the model’s confidence and its reasoning, so a reviewer can see why a file was tagged. An answer that does not fit your taxonomy, or a failed call, is stored as a fallback with zero confidence, so it is never mistaken for the model’s view.
Does the model decide where files go?
No. The model tags files. People choose each wave’s destination, and two different people sign it. Transfers, signatures and evidence do not depend on the model.
The deal’s stage decides what the platform allows.
A deal moves through nine stages, and each change of stage carries the signatures the default gate profile requires. No wave can start before Integration: the wave planner asks the deal’s stage first, and refuses if it cannot get an answer. So antitrust separation is a control rather than a policy people promise to follow.
| Stage | What the platform allows | Signatures to enter |
|---|---|---|
| Pre-deal | Discovery and classification. No wave can start. | Created by the deal team. |
| Diligence | Discovery continues. No wave can start. | One signature: the deal lead or corporate-development lead. |
| Sign to close | Discovery continues. Nothing moves to the acquirer. | Two signatures: the deal lead or corporate-development lead, and antitrust counsel. |
| Day 1 | Clean-team access ends automatically. Waves start at Integration. | Two signatures: the deal lead or corporate-development lead, and the IMO lead. |
| Integration | Wave migrations begin. | No signature: automatic, or started by the IMO lead. |
| TSA | Waves continue until the services agreement ends. | One signature: the IMO lead. |
| Closed | No new wave can start. The records stay. | Two signatures: the IMO lead, and the CIO or a tenant administrator. |
| Archived | Kept for its records. The ledger can still be re-verified. | No signature. |
| Terminated | Every deal-team assignment ends. No new wave can start. | From any stage, two signatures: the deal lead or corporate-development lead, and antitrust counsel. |
Pre-deal
Discovery and classification. No wave can start.
Created by the deal team.
Diligence
Discovery continues. No wave can start.
One signature: the deal lead or corporate-development lead.
Sign to close
Discovery continues. Nothing moves to the acquirer.
Two signatures: the deal lead or corporate-development lead, and antitrust counsel.
Day 1
Clean-team access ends automatically. Waves start at Integration.
Two signatures: the deal lead or corporate-development lead, and the IMO lead.
Integration
Wave migrations begin.
No signature: automatic, or started by the IMO lead.
TSA
Waves continue until the services agreement ends.
One signature: the IMO lead.
Closed
No new wave can start. The records stay.
Two signatures: the IMO lead, and the CIO or a tenant administrator.
Archived
Kept for its records. The ledger can still be re-verified.
No signature.
Terminated
Every deal-team assignment ends. No new wave can start.
From any stage, two signatures: the deal lead or corporate-development lead, and antitrust counsel.
Regulated signature slots, such as antitrust counsel, cannot be filled by an administrator standing in for the required signer.
Deal stage
- Pre-deal
- Diligence
- Sign to close
- Day 1
- Integration
- TSA
- Closed
Antitrust gate. Before Integration, the wave planner refuses to start a wave.
Wave 3 approval
Approved- First signerFresh login · meaning recorded
- Second signerFresh login · meaning recorded
The wave’s creator cannot be one of the two signers.
One wave: approve, execute, verify, report.
- 1
Plan
Choose the source, the destination and the file scope. A draft can be edited until it is submitted.
- 2
Sign
Two different people sign, each after a fresh login. The wave’s creator cannot approve it.
- 3
Transfer
The Go worker moves the bytes in chunks and resumes from the last good chunk after a crash.
- 4
Verify
Every file is re-read at the destination and hashed again. Any mismatch fails the wave.
- 5
Close
Closing requires the compliance report. Cancel and reopen are e-signed. Each step is written to the ledger.
Wave states
- A wave can start only in the Integration or TSA stage. The planner asks the deal’s state machine first and refuses if it cannot get an answer.
- Signatures use a fresh RS256 token from your Keycloak, checked against its published keys and valid for 300 seconds.
- Two distinct signers are enforced by the service and by a unique index in the database.
- A failed wave rolls back what it wrote at the destination, and each failed file raises an exception for a steward.
Re-read at the destination, hashed again.
A checksum the sender calculates only proves the sender did its arithmetic. Mergiva reads each file back out of the destination and hashes it again, and a wave never reports success over an unverified file.
Source system
The estate you are moving
Transfer worker
Chunked copy
Destination system
The acquirer’s target
Hash at discovery
SHA-256 of the bytes as read
Compare
The two digests
Re-read and hash again
At the destination
Match
A file-verified entry is written to the ledger with both digests.
Mismatch
The file fails, the wave is held at FAILED, and an exception is raised.
1. Hash at discovery
SHA-256 of the source bytes as they are read.
2. Transfer
The worker copies the file to the destination in chunks.
3. Re-read and hash again
The file is read back out of the destination and hashed.
4. Compare
The two digests must match.
Match
A file-verified entry is written to the ledger with both digests.
Mismatch
The file fails, the wave is held at FAILED, and an exception is raised.
Ledger
verifyChain: INTACT- 80wave.approved7e2d…b41c
- 81wave.startedprev ✓ · a90f…22e8
- 82file.verifiedprev ✓ · 3f9a…c1e7
- 83wave.completedprev ✓ · d415…7b02
Each entry carries the hash of the one before it. Database triggers reject update, delete and truncate.
Redundant, obsolete and trivial files, before you pay to move them.
Waste detection is an optional module. It reports; it never deletes.
Redundant
Identical SHA-256 within a deal. The oldest copy is kept as the original.
Obsolete
Untouched for longer than a threshold, by default three years.
Trivial
Empty files, known junk names, and non-documents under 1024 bytes.
- It reports and never deletes. The service has no delete route.
- Files under legal hold or a retention policy are shown but never counted as reclaimable.
- A file that is both redundant and trivial is counted once in the reclaimable total.
- Thresholds can be changed per installation and per scan through the API.
Run it day to day.
Reports
A Data Estate Report as a PDF from real scan data, and a compliance report when a wave closes. Reports can be scheduled daily, weekly or monthly and e-mailed through your mail relay.
Notifications
E-mail through your SMTP relay; Slack and Microsoft Teams through incoming webhooks. Every notification is logged with its delivery status.
Exceptions
A failed transfer raises an exception on its own. Stewards retry, skip or accept it, with a written reason. An AI suggestion is advisory and limited to the roles that triage; a person decides.
Retention
Policies with a fifteen-year default, applied to a wave’s verified files in one step, and per-file legal hold with a reason and an accountable user. Managed from its own screen or the API.
Sixteen built-in roles in four tiers.
The roles are built in. An installation uses all sixteen or a smaller profile, and there is no runtime role editor, so a change to a control goes through your change control.
Platform
Deal leadership
Operations
Restricted
IMO lead: the lead of the integration management office.
Start with one deal.
A pilot starts with the two systems you need to connect, and ends with evidence you can hand to an assessor.
- 1
Name the pair
Tell us the two systems you need to connect. We produce that pair’s evidence before the pilot starts.
- 2
Scan one estate
Run a Data Estate Scan in your own cluster. You get the PDF report and a classification your QA team can inspect.
- 3
Plan validation together
Evidence maps, the control inventory and test artefacts, executed with your QA team on your infrastructure.
- 4
Run the first wave
Two signatures, a verified transfer and a compliance report you can hand to an assessor.
Or write to contact@mergiva-ai.com.