Skip to content

Know the estate. Control every move. Prove every file.

Mergiva discovers and classifies an acquired estate, gates every move on the deal’s stage and two signatures, and re-hashes every file at the destination.

See what the estate holds before anything moves.

Discovery tells you what exists. Classification tells you what it is. Together they turn an unknown estate into a list of files you can scope, sign for and move in batches.

Discover

  • Scans Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online, SFTP and local or NAS sources.
  • Records each file’s SHA-256, size, content type and last-modified date.
  • Feeds the Data Estate Report, a PDF built from the scan’s real data.
  • A scan that fails shows its reason, so your team can fix it and run it again.
  • Allowed from pre-deal through the TSA period, so diligence can start before signing.

Classify

  • AI-assisted by default. Microsoft Presidio redacts detected personal data before a sample leaves your installation.
  • You choose the model for each AI feature, and whether calls go to Anthropic with your key or through your own AI gateway.
  • A wave’s file scope can be narrowed by classification category and sensitivity, as well as by path, file type and size.

Data Estate Scan

Source: SFTP
FileCategorySens.
CSR_0142_final.pdfGxP-Clinical4
patient_listing_v3.xlsxPHI5
batch_record_BR-2231.pdfGxP-Manufacturing4
HSR_filing_notes.docxLegal-Antitrust5
FY25_forecast.xlsxFinancial3
site_photos.zipOther1
1 Public2 Internal3 Confidential4 Restricted5 Highly Restricted

AI classification, under your control.

Mergiva uses a large language model to place every discovered file in your taxonomy and score its sensitivity. You choose the model for each AI feature, the account it runs under and how personal data is handled.

In a life-sciences deal the first questions are not about bytes. Where are the clinical records? Which files hold personal or health data? What is competitively sensitive, and must stay behind the antitrust gate until Day 1? The answers decide how the estate is split into waves, who signs for each one, and what your privacy and quality teams need to see.

Mergiva answers them file by file. Once discovery has recorded and hashed every file, a large language model reads each file’s details, places it in your taxonomy and scores its sensitivity from 1 to 5. Every answer is stored with the model’s confidence and its reasoning, so a reviewer can see why a file was tagged the way it was.

You stay in charge of the AI. You choose the model and the account it runs under, how personal data is handled before a request leaves your cluster, and how much file content the model may see. If your policy is that no personal data leaves, set the guard to refuse any classification request that contains it.

How one file is classified

Six steps, and five of them run inside your cluster.

  1. Step 1

    Discovered file

    Name, path and type

  2. Step 2

    Prompt from your taxonomy

    Your categories and scale

  3. Step 3

    Personal data guard

    Redact, block or warn

  4. Step 4The one outside call

    Your chosen model

    Claude, or your gateway

  5. Step 5

    Checked answer

    Only your tags accepted

  6. Step 6

    Stored with the file

    Tags, score, confidence

Steps 1 to 3 and 5 to 6 run inside your cluster. Step 4 is the only one that leaves it, and only after the personal-data guard.
  1. 1

    Discovered file

    The discovery scan has already recorded each file’s name, folder path, content type and SHA-256. For files on a local or NAS source, the classifier also reads up to the first 4 KB.

  2. 2

    Prompt from your taxonomy

    The classifier builds the prompt from your taxonomy profile: every category with its definition, and the sensitivity scale. It asks for one answer in a fixed format.

  3. 3

    Personal data guard

    Before the request leaves your cluster, Microsoft Presidio looks for personal data. By default it redacts what it finds. You can make it refuse the call instead, or only log the finding.

  4. 4

    Your chosen model

    The request goes to the model you chose for classification, Claude Haiku 4.5 unless you set another. It travels to Anthropic under your own key, or through your own AI gateway.

  5. 5

    Checked answer

    The answer is parsed and checked. Tags outside your taxonomy are dropped, and the score must sit on your scale. If the answer fails the check or the call fails, the file is marked as a fallback with zero confidence, so no one mistakes it for the model’s view.

  6. 6

    Stored with the file

    The result is stored against the file: its category tags, a sensitivity score, the model’s confidence and reasoning, and whether it came from the model or the fallback.

Exactly what goes into the prompt

What the model receives depends on where the file lives. The prompt holds nothing else: no deal name, no hashes, no other files.

Local folders and NAS shares
File name, folder path, content type, and up to the first 4 KB of the file. Text files are sent as text; other files as a short hexadecimal excerpt.
Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP
File name, folder path and content type. No file content.
Every request
Your taxonomy’s categories with their definitions, the 1 to 5 sensitivity scale, and the format the answer must take. By default, detected personal data is redacted from the file’s details before the call.

Classification requestLeaves your cluster

file name
patient_listing_v3.xlsx
path
/clinical/site-07/
type
spreadsheet
sample
Patient [REDACTED], visit 3 …
taxonomy
8 categories, scale 1 to 5

The sample is sent for local and NAS files only, up to 4 KB.

Answer, stored with the file

PHIGxP-Clinical
sensitivity
5 · Highly Restricted
confidence
0.94
reasoning
patient-level clinical listing
Illustration with sample data.

Your model, your account, your rules

Each setting below exists in the product today. The note under each says where it is set.

Your model, feature by feature
Choose the model for each AI feature, classification included. Claude Haiku 4.5 is the default, and Claude Sonnet 4.6 and Claude Opus 4.6 are also offered. Behind your own gateway, you list the models it serves.
AI Operations screen, by an administrator
Your own account, or your own gateway
Calls run under the API key you give the installation, so they fall under your own agreement with Anthropic. Or point Mergiva at your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. Every call goes only where you route it. A call your settings cannot serve is refused, never sent somewhere else.
Installation settings
How personal data is handled
Redact detected personal data before the call (the default), refuse any request that contains it, or only log it. A second, optional check scans the model’s answer and redacts any personal data it repeats.
Installation settings
Your taxonomy
The categories, their definitions and the names of the five sensitivity levels come from a taxonomy profile. The default is the pharma profile shown on this page. You can supply your own; the scale stays 1 to 5.
A profile file, read when the service starts
How much content the model sees
Set how much of a local or NAS file goes into the prompt, from the default of 4 KB down to none.
Installation settings
Rate limits and spend budgets
Cap requests and tokens per minute, concurrent calls and the size of a single request. Set daily or monthly spend budgets for each tenant. Both are checked before every classification call.
By a platform or tenant administrator
A circuit breaker
When the model keeps failing, Mergiva stops calling it until it recovers, so an outage fails fast instead of stalling every scan. Calls refused for rate limiting are retried with a growing delay.
Built in
Usage you can see
The AI Operations screen shows calls, tokens, cost, cache hits and latency for classification, with the model each call used.
AI Operations screen

Eight categories and a five-level scale

The categories and definitions the model is given, exactly as the default pharma profile states them. Replace them with your own profile if your organisation classifies differently.

GxP-Clinical
clinical trial data, protocols, case report forms
GxP-Manufacturing
batch records, deviations, equipment validation
GxP-Regulatory
FDA/EMA submissions, 510(k), CMC, CTA, labeling
PHI
personal health information (named patients, MRNs, diagnoses)
PII
personal information without health context (HR records, contractor names)
Financial
revenue, accounting, internal financials
Legal-Antitrust
HSR, FTC/DOJ filings, antitrust counsel work product
Other
none of the above

Sensitivity, scored 1 to 5

  1. 1Public
  2. 2Internal
  3. 3Confidential
  4. 4Restricted
  5. 5Highly Restricted

A file can carry more than one category, such as clinical and health data together. It always carries exactly one sensitivity score.

Classification shapes the rest of the deal

Scoping waves

A wave’s file scope can be narrowed by category and sensitivity, so clinical records, personal data or financial files can move in their own approved batch.

The Data Estate Report

A Data Estate Scan runs discovery, then classification, then builds the PDF report, so the deal team sees what the estate holds before anyone plans a move.

The Classification Explorer

Reviewers filter and browse the results, and open any file to see its tags, score, confidence and the model’s reasoning.

What buyers ask about the AI

Which AI model does Mergiva use?

Claude Haiku 4.5 by default. Claude Sonnet 4.6 and Claude Opus 4.6 are also supported, and an administrator can choose the model for each AI feature. Behind your own gateway, you list the models it serves.

Can we use our own AI account?

Yes. The installation calls the model with the API key you give it, under your own agreement with Anthropic. Or route every call through your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. A call your settings cannot serve is refused, never sent somewhere else.

What exactly does the model see?

For files on local or NAS sources: the name, folder path, content type and up to the first 4 KB. For Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP: the name, path and type only. By default, detected personal data is redacted before the call.

How can we check what the model decided?

Each result carries the model’s confidence and its reasoning, so a reviewer can see why a file was tagged. An answer that does not fit your taxonomy, or a failed call, is stored as a fallback with zero confidence, so it is never mistaken for the model’s view.

Does the model decide where files go?

No. The model tags files. People choose each wave’s destination, and two different people sign it. Transfers, signatures and evidence do not depend on the model.

The deal’s stage decides what the platform allows.

A deal moves through nine stages, and each change of stage carries the signatures the default gate profile requires. No wave can start before Integration: the wave planner asks the deal’s stage first, and refuses if it cannot get an answer. So antitrust separation is a control rather than a policy people promise to follow.

  • Pre-deal

    Discovery and classification. No wave can start.

    Created by the deal team.

  • Diligence

    Discovery continues. No wave can start.

    One signature: the deal lead or corporate-development lead.

  • Antitrust gate Sign to close

    Discovery continues. Nothing moves to the acquirer.

    Two signatures: the deal lead or corporate-development lead, and antitrust counsel.

  • Day 1

    Clean-team access ends automatically. Waves start at Integration.

    Two signatures: the deal lead or corporate-development lead, and the IMO lead.

  • Integration

    Wave migrations begin.

    No signature: automatic, or started by the IMO lead.

  • TSA

    Waves continue until the services agreement ends.

    One signature: the IMO lead.

  • Closed

    No new wave can start. The records stay.

    Two signatures: the IMO lead, and the CIO or a tenant administrator.

  • Archived

    Kept for its records. The ledger can still be re-verified.

    No signature.

  • Terminated

    Every deal-team assignment ends. No new wave can start.

    From any stage, two signatures: the deal lead or corporate-development lead, and antitrust counsel.

Regulated signature slots, such as antitrust counsel, cannot be filled by an administrator standing in for the required signer.

Deal stage

  1. Pre-deal
  2. Diligence
  3. Sign to close
  4. Day 1
  5. Integration
  6. TSA
  7. Closed

Antitrust gate. Before Integration, the wave planner refuses to start a wave.

Wave 3 approval

Approved
  • First signerFresh login · meaning recorded
  • Second signerFresh login · meaning recorded

The wave’s creator cannot be one of the two signers.

One wave: approve, execute, verify, report.

  1. 1

    Plan

    Choose the source, the destination and the file scope. A draft can be edited until it is submitted.

  2. 2

    Sign

    Two different people sign, each after a fresh login. The wave’s creator cannot approve it.

  3. 3

    Transfer

    The Go worker moves the bytes in chunks and resumes from the last good chunk after a crash.

  4. 4

    Verify

    Every file is re-read at the destination and hashed again. Any mismatch fails the wave.

  5. 5

    Close

    Closing requires the compliance report. Cancel and reopen are e-signed. Each step is written to the ledger.

Wave states

DraftPending approvalApprovedRunning / PausedCompletedClosedFailedRolled backCancelled
  • A wave can start only in the Integration or TSA stage. The planner asks the deal’s state machine first and refuses if it cannot get an answer.
  • Signatures use a fresh RS256 token from your Keycloak, checked against its published keys and valid for 300 seconds.
  • Two distinct signers are enforced by the service and by a unique index in the database.
  • A failed wave rolls back what it wrote at the destination, and each failed file raises an exception for a steward.

Re-read at the destination, hashed again.

A checksum the sender calculates only proves the sender did its arithmetic. Mergiva reads each file back out of the destination and hashes it again, and a wave never reports success over an unverified file.

  1. 1. Hash at discovery

    SHA-256 of the source bytes as they are read.

  2. 2. Transfer

    The worker copies the file to the destination in chunks.

  3. 3. Re-read and hash again

    The file is read back out of the destination and hashed.

  4. 4. Compare

    The two digests must match.

  5. Match

    A file-verified entry is written to the ledger with both digests.

    Mismatch

    The file fails, the wave is held at FAILED, and an exception is raised.

Ledger

verifyChain: INTACT
  1. 80
    wave.approved7e2d…b41c
  2. 81
    wave.starteda90f…22e8
  3. 82
    file.verified3f9a…c1e7
  4. 83
    wave.completedd415…7b02

Each entry carries the hash of the one before it. Database triggers reject update, delete and truncate.

Redundant, obsolete and trivial files, before you pay to move them.

Waste detection is an optional module. It reports; it never deletes.

Redundant

Identical SHA-256 within a deal. The oldest copy is kept as the original.

Obsolete

Untouched for longer than a threshold, by default three years.

Trivial

Empty files, known junk names, and non-documents under 1024 bytes.

  • It reports and never deletes. The service has no delete route.
  • Files under legal hold or a retention policy are shown but never counted as reclaimable.
  • A file that is both redundant and trivial is counted once in the reclaimable total.
  • Thresholds can be changed per installation and per scan through the API.

Run it day to day.

Reports

A Data Estate Report as a PDF from real scan data, and a compliance report when a wave closes. Reports can be scheduled daily, weekly or monthly and e-mailed through your mail relay.

Notifications

E-mail through your SMTP relay; Slack and Microsoft Teams through incoming webhooks. Every notification is logged with its delivery status.

Exceptions

A failed transfer raises an exception on its own. Stewards retry, skip or accept it, with a written reason. An AI suggestion is advisory and limited to the roles that triage; a person decides.

Retention

Policies with a fifteen-year default, applied to a wave’s verified files in one step, and per-file legal hold with a reason and an accountable user. Managed from its own screen or the API.

Sixteen built-in roles in four tiers.

The roles are built in. An installation uses all sixteen or a smaller profile, and there is no runtime role editor, so a change to a control goes through your change control.

Platform

Platform adminTenant admin

Deal leadership

Deal ownerIMO leadDeal leadCorporate development leadWorkstream lead

Operations

StewardReviewerOperator

Restricted

Clean teamObserverAcquired employeeBuyer representativeAntitrust counselCIO

IMO lead: the lead of the integration management office.

Start with one deal.

A pilot starts with the two systems you need to connect, and ends with evidence you can hand to an assessor.

  1. 1

    Name the pair

    Tell us the two systems you need to connect. We produce that pair’s evidence before the pilot starts.

  2. 2

    Scan one estate

    Run a Data Estate Scan in your own cluster. You get the PDF report and a classification your QA team can inspect.

  3. 3

    Plan validation together

    Evidence maps, the control inventory and test artefacts, executed with your QA team on your infrastructure.

  4. 4

    Run the first wave

    Two signatures, a verified transfer and a compliance report you can hand to an assessor.

Or write to contact@mergiva-ai.com.