Skip to content

AI classification

Know what every file is before anything moves

Mergiva uses a large language model to place every discovered file in your taxonomy and score its sensitivity. You choose the model for each AI feature, the account it runs under and how personal data is handled.

  • Claude Haiku 4.5 by default
  • Your key or your gateway
  • Personal data redacted first
  • Model chosen per feature

Why classify first

The plan for a regulated estate starts with what is in it

In a life-sciences deal the first questions are not about bytes. Where are the clinical records? Which files hold personal or health data? What is competitively sensitive, and must stay behind the antitrust gate until Day 1? The answers decide how the estate is split into waves, who signs for each one, and what your privacy and quality teams need to see.

Mergiva answers them file by file. Once discovery has recorded and hashed every file, a large language model reads each file’s details, places it in your taxonomy and scores its sensitivity from 1 to 5. Every answer is stored with the model’s confidence and its reasoning, so a reviewer can see why a file was tagged the way it was.

You stay in charge of the AI. You choose the model and the account it runs under, how personal data is handled before a request leaves your cluster, and how much file content the model may see. If your policy is that no personal data leaves, set the guard to refuse any classification request that contains it.

How it works

How one file is classified

Six steps, and five of them run inside your cluster. The diagram shows the path a single file takes; the notes below say what happens at each step.

  1. Step 1

    Discovered file

    Name, path and type

  2. Step 2

    Prompt from your taxonomy

    Your categories and scale

  3. Step 3

    Personal data guard

    Redact, block or warn

  4. Step 4The one outside call

    Your chosen model

    Claude, or your gateway

  5. Step 5

    Checked answer

    Only your tags accepted

  6. Step 6

    Stored with the file

    Tags, score, confidence

Steps 1 to 3 and 5 to 6 run inside your cluster. Step 4 is the only one that leaves it, and only after the personal-data guard.
  1. 1

    Discovered file

    The discovery scan has already recorded each file’s name, folder path, content type and SHA-256. For files on a local or NAS source, the classifier also reads up to the first 4 KB.

  2. 2

    Prompt from your taxonomy

    The classifier builds the prompt from your taxonomy profile: every category with its definition, and the sensitivity scale. It asks for one answer in a fixed format.

  3. 3

    Personal data guard

    Before the request leaves your cluster, Microsoft Presidio looks for personal data. By default it redacts what it finds. You can make it refuse the call instead, or only log the finding.

  4. 4

    Your chosen model

    The request goes to the model you chose for classification, Claude Haiku 4.5 unless you set another. It travels to Anthropic under your own key, or through your own AI gateway.

  5. 5

    Checked answer

    The answer is parsed and checked. Tags outside your taxonomy are dropped, and the score must sit on your scale. If the answer fails the check or the call fails, the file is marked as a fallback with zero confidence, so no one mistakes it for the model’s view.

  6. 6

    Stored with the file

    The result is stored against the file: its category tags, a sensitivity score, the model’s confidence and reasoning, and whether it came from the model or the fallback.

What the model sees

Exactly what goes into the prompt

What the model receives depends on where the file lives. The prompt holds nothing else: no deal name, no hashes, no other files.

Local folders and NAS shares
File name, folder path, content type, and up to the first 4 KB of the file. Text files are sent as text; other files as a short hexadecimal excerpt.
Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP
File name, folder path and content type. No file content.
Every request
Your taxonomy’s categories with their definitions, the 1 to 5 sensitivity scale, and the format the answer must take. By default, detected personal data is redacted from the file’s details before the call.

You configure it

Your model, your account, your rules

Each setting below exists in the product today. The label under each one says where it is set.

  • Your model, feature by feature

    Choose the model for each AI feature, classification included. Claude Haiku 4.5 is the default, and Claude Sonnet 4.6 and Claude Opus 4.6 are also offered. Behind your own gateway, you list the models it serves.

    Where: AI Operations screen, by an administrator

  • Your own account, or your own gateway

    Calls run under the API key you give the installation, so they fall under your own agreement with Anthropic. Or point Mergiva at your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. Every call goes only where you route it. A call your settings cannot serve is refused, never sent somewhere else.

    Where: Installation settings

  • How personal data is handled

    Redact detected personal data before the call (the default), refuse any request that contains it, or only log it. A second, optional check scans the model’s answer and redacts any personal data it repeats.

    Where: Installation settings

  • Your taxonomy

    The categories, their definitions and the names of the five sensitivity levels come from a taxonomy profile. The default is the pharma profile shown on this page. You can supply your own; the scale stays 1 to 5.

    Where: A profile file, read when the service starts

  • How much content the model sees

    Set how much of a local or NAS file goes into the prompt, from the default of 4 KB down to none.

    Where: Installation settings

  • Rate limits and spend budgets

    Cap requests and tokens per minute, concurrent calls and the size of a single request. Set daily or monthly spend budgets for each tenant. Both are checked before every classification call.

    Where: By a platform or tenant administrator

  • A circuit breaker

    When the model keeps failing, Mergiva stops calling it until it recovers, so an outage fails fast instead of stalling every scan. Calls refused for rate limiting are retried with a growing delay.

    Where: Built in

  • Usage you can see

    The AI Operations screen shows calls, tokens, cost, cache hits and latency for classification, with the model each call used.

    Where: AI Operations screen

The default taxonomy

Eight categories and a five-level scale

These are the categories and definitions the model is given, exactly as the default pharma profile states them. Replace them with your own profile if your organisation classifies differently.

GxP-Clinical
clinical trial data, protocols, case report forms
GxP-Manufacturing
batch records, deviations, equipment validation
GxP-Regulatory
FDA/EMA submissions, 510(k), CMC, CTA, labeling
PHI
personal health information (named patients, MRNs, diagnoses)
PII
personal information without health context (HR records, contractor names)
Financial
revenue, accounting, internal financials
Legal-Antitrust
HSR, FTC/DOJ filings, antitrust counsel work product
Other
none of the above

Sensitivity, scored 1 to 5

  1. 1Public
  2. 2Internal
  3. 3Confidential
  4. 4Restricted
  5. 5Highly Restricted

A file can carry more than one category, such as clinical and health data together. It always carries exactly one sensitivity score.

Where results go

Classification shapes the rest of the deal

Scoping waves

A wave’s file scope can be narrowed by category and sensitivity, so clinical records, personal data or financial files can move in their own approved batch.

The Data Estate Report

A Data Estate Scan runs discovery, then classification, then builds the PDF report, so the deal team sees what the estate holds before anyone plans a move.

The Classification Explorer

Reviewers filter and browse the results, and open any file to see its tags, score, confidence and the model’s reasoning.

Questions

What buyers ask about the AI

Which AI model does Mergiva use?

Claude Haiku 4.5 by default. Claude Sonnet 4.6 and Claude Opus 4.6 are also supported, and an administrator can choose the model for each AI feature. Behind your own gateway, you list the models it serves.

Can we use our own AI account?

Yes. The installation calls the model with the API key you give it, under your own agreement with Anthropic. Or route every call through your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. A call your settings cannot serve is refused, never sent somewhere else.

What exactly does the model see?

For files on local or NAS sources: the name, folder path, content type and up to the first 4 KB. For Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP: the name, path and type only. By default, detected personal data is redacted before the call.

How can we check what the model decided?

Each result carries the model’s confidence and its reasoning, so a reviewer can see why a file was tagged. An answer that does not fit your taxonomy, or a failed call, is stored as a fallback with zero confidence, so it is never mistaken for the model’s view.

Does the model decide where files go?

No. The model tags files. People choose each wave’s destination, and two different people sign it. Transfers, signatures and evidence do not depend on the model.

Start with one deal.

Judge us on the ledger, not the demo.

Talk to us
  1. 1

    Name the pair

    Tell us the two systems you need to connect. We produce that pair’s evidence before the pilot starts.

  2. 2

    Scan one estate

    Run a Data Estate Scan in your own cluster. You get the PDF report and a classification your QA team can inspect.

  3. 3

    Plan validation together

    Evidence maps, the control inventory and test artefacts, executed with your QA team on your infrastructure.

  4. 4

    Run the first wave

    Two signatures, a verified transfer and a compliance report you can hand to an assessor.

Or write to contact@mergiva-ai.com.