Nothing is sent, and you can check it
Klarmask does not ask you to take its word for it. Two promises, and the means to check each one: the document does not leave your computer, and what is redacted no longer exists in the file.
Nothing is sent
The document is read by your browser, in the open tab. Detection, review, rebuilding the file and checking it all happen there too. There is no processing server: there is nowhere to send the document.
The only requests the page makes are reads from this site: the pages, the code, the standard fonts of the PDF format and, if you ask for it, the name model, once, verified by its fingerprint then stored on your computer. No account, no cookie, no analytics. The only request that is not a read is the counter, described below. The server keeps no access log.
An architecture rather than a promise
An online service can promise to delete your document; you have no way of seeing that it does. An architecture, on the other hand, can be seen.
| An online redaction service | Klarmask | |
|---|---|---|
| Where the document is processed | On the provider's servers, or at its artificial intelligence supplier | In your browser, on your computer |
| What you have to believe | Its retention policy and its terms | Nothing: cut the network and watch |
| Processors | Host, model supplier, sometimes outside the European Union | None for the document: there is no processing to subcontract |
| Retention period | Whatever its terms say, to be negotiated | Not applicable: nothing is received |
| Reuse to train a model | Depending on its terms | Impossible: we receive nothing |
| Dependence | On the service, its prices, the law of its country | None: the self-hosted version runs on your network |
| Account, cookies | Often required | None |
Check it yourself
Cut the network
Load the page, then turn off Wi-Fi or unplug the cable. Drop a document, decide, produce the file: everything goes through to the end.
Watch the Network tab
Open the developer tools (F12), “Network” tab, before dropping the document. While you work, no line is added; when you save the file, a single one,
/compte, the counter described below. The workspace also shows its own list, below the review; but a page could lie about itself, the browser's tab cannot.Read the security policy
Every page carries a content security policy that the browser enforces. The line that matters:
connect-src 'self'. It forbids any connection to another site, even if the page's code asked for one.
The only thing that leaves: a counter
When you save a redacted file, or when a document has been checked, your browser sends one request to this site, and only one: POST /compte?f=caviardage&p=12. It says whether it was a redaction or a check, and how many pages the document had. Nothing else: not a word of the document, not a redacted value, not a file name, not an identifier.
The server writes only one line: the date, the category, the number of pages. No IP address, no browser, no referring page: two documents from the same person cannot be told apart from two documents from strangers. The total is published as is on the home page.
A browser that asks not to be tracked (Do Not Track, Global Privacy Control) sends nothing. Offline, nothing leaves, and nothing is kept for later. The self-hosted version is built without the counter. In every case, the tool works exactly the same.
To put it plainly: we record the number of documents redacted and checked, and their number of pages. Nothing in these lines says who processed them, from where, or what they contained.
What this changes for the GDPR
We receive no document, so we process none of the data it contains. In practice, for an organisation:
- No transfer of data to a third party: the document does not leave the computer of the person redacting it. The question of transfers outside the European Union does not arise, since there is no transfer.
- No processor to qualify for the content of the documents, so no data processing agreement to sign for the redaction itself.
- No retention period to set and no deletion to request: there is no copy anywhere else. The open document disappears when you close the tab.
- No artificial intelligence supplier sees your documents: the name model is published under an open licence (MIT), pinned, verified and served by this site, then it runs on your computer.
- No cookie, no tracker, no third party: no consent banner, because there is nothing to consent to. The site's only measurement is the counter described above.
- The site is hosted in France, by OVH, and serves only files. To depend on nobody, the self-hosted version runs on your internal network.
What remains your responsibility
Klarmask does not decide for you what must be redacted, and does not guarantee that a redacted document no longer allows someone to be recognised: your reread does that. If you pseudonymise so that you can re-identify later, the mapping between labels and people remains the additional information within the meaning of the GDPR: keep it separately. The terms of use say so too.
What is stored on your computer
The name model and its engine, if you download them, are stored in the private space the browser reserves for this site, after their fingerprint has been verified. This avoids downloading them again. The workspace lets you delete them with one button.
Your document, however, is stored nowhere: it is read into memory, in the tab, and disappears when you close it. The file produced is saved by your browser, wherever you decide.
What is redacted no longer exists in the file
The file produced is a new document. The original document is never edited then saved again: that is how redacted documents have been unredacted, through the previous version left in the file.
A Word document produced is also a new package, written from only the parts the tool understands: tracked changes accepted, comments, properties, charts and embedded objects removed, image metadata erased.
| Where data can hide | What Klarmask does |
|---|---|
| Text under a black rectangle | The page is rendered as an image and the black box is painted onto the pixels. The original text is copied nowhere. |
| Searchable text | Rewritten from the kept words only. A word with a single redacted letter is not included. |
| White-on-white, off-page or tiny text | Read by detection, so suggested in the review; erased by rendering as an image. |
| Annotations, comments, form fields | Not copied. |
| Bookmarks, attachments, JavaScript | Not copied. |
| Metadata and XMP | Only the software name remains. No author, no title, no date. |
| Earlier versions | The file produced has a single version. |
| The file name | A neutral name is suggested: “redacted-document.pdf”. |
The name model
Names are suggested by a named entity recognition model, CamemBERT-NER (MIT licence), trained on French. It is downloaded only if you ask for it, from this site and not from a third party, once: each file is verified by its SHA-256 fingerprint before being stored in a space the browser reserves for this site, on your computer. It reads only the text of the open page, in the tab.
The library that runs it is not allowed onto the network: its download function is replaced by one that refuses, and it reads only files already verified. The fingerprints published here are the ones the tool checks.
Model : Xenova/camembert-ner, revision 8e1988247e87.
| File | Size | SHA-256 |
|---|---|---|
ort/ort-wasm-simd-threaded.asyncify.mjs | 0.1 MB | 0966b6105cd936744498aa60df7a22cbd47af3374dbc64a9ab561c08a71e3611 |
ort/ort-wasm-simd-threaded.asyncify.wasm | 26.9 MB | 49871f5a4409519797e127440868a6d1923339d9185907f301a5b2a1d90af082 |
modeles/Xenova/camembert-ner/config.json | 0 MB | 3f928a29f1ee4578745e0a989f2d35666df49c7b431c8a3e01f52be661279699 |
modeles/Xenova/camembert-ner/tokenizer.json | 2.4 MB | 2fe563b59de779621fa6ef3756e7ae565a87d78ac249afb728daf743ba3821ed |
modeles/Xenova/camembert-ner/tokenizer_config.json | 0 MB | ec016bffadf7b3ffc0ca3b9318d166b520a364c8ea3e40b0e26245738b1c08b5 |
modeles/Xenova/camembert-ner/special_tokens_map.json | 0 MB | 5bf5488fdfc958edafccea2aebff4c8b15e383c16bf2061e2cc955688502af8b |
modeles/Xenova/camembert-ner/onnx/model_quantized.onnx | 111.3 MB | 43507a27e7e5500b5e197c8c7de803bc94bcc29ce156fcbb19bda18ed322960d |
modeles/Xenova/camembert-ner/onnx/model.onnx | 440.4 MB | 0a2c64c0f051642a3002e5cef91861a26eaa44cceedceac4ce8eabee3fac4f75 |
The output check
Before offering you the file, the tool reads it back by paths independent of the one that wrote it: the text, with a second PDF reader; the file's objects, walked through one by one, each stream decompressed. Each redacted item is searched for in all its forms (spaces, accents, capitals). The structure is checked: no annotation, no form, no bookmark, no attachment, no metadata, and a single version.
If it finds anything at all, the file is not offered, and the screen says what was found and where.
What is not guaranteed
- Browser extensions can read the pages you open. For a sensitive document, use a profile with no extensions.
- A compromised computer (spyware, screen capture) sees everything you see. No website can do anything about it.
- The browser itself sends information to its publisher (updates, crash reports) depending on its settings.
- Context: a document whose names are redacted can remain recognisable. See pseudonymisation or anonymisation.
- Detection suggests; it does not guarantee it has seen everything. The reread is yours.