Skip to content
Klarmask

Nothing is sent, and you can check it

Klarmask does not ask you to take its word for it. Two promises, and the means to check each one: the document does not leave your computer, and what is redacted no longer exists in the file.

Nothing is sent

The document is read by your browser, in the open tab. Detection, review, rebuilding the file and checking it all happen there too. There is no processing server: there is nowhere to send the document.

The only requests the page makes are reads from this site: the pages, the code, the standard fonts of the PDF format and, if you ask for it, the name model, and, if you drop a scan or a photo, the text recognition engine and the language model, once, verified by its fingerprint then stored on your computer. No account, no cookie, no analytics. The only request that is not a read is the counter, described below. The server keeps no access log.

An architecture rather than a promise

An online service can promise to delete your document; you have no way of seeing that it does. An architecture, on the other hand, can be seen.

An online redaction serviceKlarmask
Where the document is processedOn the provider's servers, or at its artificial intelligence supplierIn your browser, on your computer
What you have to believeIts retention policy and its termsNothing: cut the network and watch
ProcessorsHost, model supplier, sometimes outside the European UnionNone for the document: there is no processing to subcontract
Retention periodWhatever its terms say, to be negotiatedNot applicable: nothing is received
Reuse to train a modelDepending on its termsImpossible: we receive nothing
DependenceOn the service, its prices, the law of its countryNone: the self-hosted version runs on your network
Account, cookiesOften requiredNone

Check it yourself

  1. Cut the network

    Load the page, then turn off Wi-Fi or unplug the cable. Drop a document, decide, produce the file: everything goes through to the end. Even reload the page: after a first visit, the tool reopens without a network.

  2. Watch the Network tab

    Open the developer tools (F12), “Network” tab, before dropping the document. While you work, no line is added; when you save the file, a single one, /compte, the counter described below. The workspace also shows its own list, below the review; but a page could lie about itself, the browser's tab cannot.

  3. Read the security policy

    Every page carries a content security policy that the browser enforces. The line that matters: connect-src 'self'. It forbids any connection to another site, even if the page's code asked for one.

The only thing that leaves: a counter

When you save a redacted file, or when a document has been checked, your browser sends one request to this site, and only one: POST /compte?f=caviardage&p=12. It says whether it was a redaction or a check, and how many pages the document had. Nothing else: not a word of the document, not a redacted value, not a file name, not an identifier.

The server writes only one line: the date, the category, the number of pages. No IP address, no browser, no referring page: two documents from the same person cannot be told apart from two documents from strangers. The total is published as is on the home page.

A browser that asks not to be tracked (Do Not Track, Global Privacy Control) sends nothing. Offline, nothing leaves, and nothing is kept for later. The self-hosted version is built without the counter. In every case, the tool works exactly the same.

To put it plainly: we record the number of documents redacted and checked, and their number of pages. Nothing in these lines says who processed them, from where, or what they contained.

What this changes for the GDPR

We receive no document, so we process none of the data it contains. In practice, for an organisation:

  • No transfer of data to a third party: the document does not leave the computer of the person redacting it. The question of transfers outside the European Union does not arise, since there is no transfer.
  • No processor to qualify for the content of the documents, so no data processing agreement to sign for the redaction itself.
  • No retention period to set and no deletion to request: there is no copy anywhere else. The open document disappears when you close the tab.
  • No artificial intelligence supplier sees your documents: the name model is published under an open licence (MIT), pinned, verified and served by this site, then it runs on your computer.
  • No cookie, no tracker, no third party: no consent banner, because there is nothing to consent to. The site's only measurement is the counter described above.
  • The site is hosted in France, by OVH, and serves only files. To depend on nobody, the self-hosted version runs on your internal network.

What remains your responsibility

Klarmask does not decide for you what must be redacted, and does not guarantee that a redacted document no longer allows someone to be recognised: your reread does that. If you pseudonymise so that you can re-identify later, the mapping between labels and people remains the additional information within the meaning of the GDPR: keep it separately. The terms of use say so too.

What is stored on your computer

The name model and its engine, if you download them, are stored in the private space the browser reserves for this site, after their fingerprint has been verified. This avoids downloading them again. The workspace lets you delete them with one button.

The pages and the code of the workspace are kept by the browser (a service worker), so that the tool reopens without a network. The scan reading engine is kept there after its first use, and is still verified by its fingerprint each time it is used. This worker makes no request of its own: it does not touch the counter and never replays anything.

If you create any, your rules (the names or words to always redact) are kept in this browser's storage, for this site: they do not leave the computer, and are erased with the site data.

Your document, however, is stored nowhere: it is read into memory, in the tab, and disappears when you close it. The file produced is saved by your browser, wherever you decide.

What is redacted no longer exists in the file

The file produced is a new document. The original document is never edited then saved again: that is how redacted documents have been unredacted, through the previous version left in the file.

A Word document produced is also a new package, written from only the parts the tool understands: tracked changes accepted, comments, properties, charts and embedded objects removed, image metadata erased.

Where data can hideWhat Klarmask does
Text under a black rectangleThe page is rendered as an image and the black box is painted onto the pixels. The original text is copied nowhere.
Searchable textRewritten from the kept words only. A word with a single redacted letter is not included.
White-on-white, off-page or tiny textRead by detection, so suggested in the review; erased by rendering as an image.
Annotations, comments, form fieldsNot copied.
Bookmarks, attachments, JavaScriptNot copied.
Metadata and XMPOnly the software name remains. No author, no title, no date.
Earlier versionsThe file produced has a single version.
The file nameA neutral name is suggested: “redacted-document.pdf”.

Reading scans

A page with no text is read by Tesseract (Apache 2.0 licence), the Klarfile engine, compiled for the browser. The engine and the language model come from this site, not from a content delivery network: the usual library that wraps them would fetch them elsewhere, so it is not used. Each file is verified by its fingerprint before use; the fingerprints are published below, with those of the name model.

Reading happens in a separate thread, in the tab. The page image does not leave it.

The name model

Names are suggested by a named entity recognition model, CamemBERT-NER (MIT licence), trained on French. It is downloaded only if you ask for it, from this site and not from a third party, once: each file is verified by its SHA-256 fingerprint before being stored in a space the browser reserves for this site, on your computer. It reads only the text of the open page, in the tab.

The library that runs it is not allowed onto the network: its download function is replaced by one that refuses, and it reads only files already verified. The fingerprints published here are the ones the tool checks.

Model : Xenova/camembert-ner, revision 8e1988247e87.

FileSizeSHA-256
ort/ort-wasm-simd-threaded.asyncify.mjs0.1 MB0966b6105cd936744498aa60df7a22cbd47af3374dbc64a9ab561c08a71e3611
ort/ort-wasm-simd-threaded.asyncify.wasm26.9 MB49871f5a4409519797e127440868a6d1923339d9185907f301a5b2a1d90af082
modeles/Xenova/camembert-ner/config.json0 MB3f928a29f1ee4578745e0a989f2d35666df49c7b431c8a3e01f52be661279699
modeles/Xenova/camembert-ner/tokenizer.json2.4 MB2fe563b59de779621fa6ef3756e7ae565a87d78ac249afb728daf743ba3821ed
modeles/Xenova/camembert-ner/tokenizer_config.json0 MBec016bffadf7b3ffc0ca3b9318d166b520a364c8ea3e40b0e26245738b1c08b5
modeles/Xenova/camembert-ner/special_tokens_map.json0 MB5bf5488fdfc958edafccea2aebff4c8b15e383c16bf2061e2cc955688502af8b
modeles/Xenova/camembert-ner/onnx/model_quantized.onnx111.3 MB43507a27e7e5500b5e197c8c7de803bc94bcc29ce156fcbb19bda18ed322960d
modeles/Xenova/camembert-ner/onnx/model.onnx440.4 MB0a2c64c0f051642a3002e5cef91861a26eaa44cceedceac4ce8eabee3fac4f75
ocr/tesseract-core-simd-lstm.js0.1 MBbe3504705d7111d1d1f3f7f9dff326c26d334031ede36e31c4d3cf883027e982
ocr/tesseract-core-simd-lstm.wasm2.9 MB187d76742dfc0d8929f0b49a619f145bb6370730776c7bd0d3e20c6b2098808d
ocr/tesseract-core-lstm.js0.1 MB48a3ee8e00924cb8c7f0cc0d099b1318fea120af56b3ee8fb3a70dd2311806c2
ocr/tesseract-core-lstm.wasm2.9 MB220e2e87551edccb85519796a170469f8ab2a8055216789e3b8b1ada18b7bc2b
ocr/fra.traineddata.gz0.7 MBd611139672b3752c7097e671e4a1d9209dfd37f2aeb081ef6487fba3351e9255
ocr/eng.traineddata.gz3 MB45b4cb346724ac1774f1c36f42f182b887bcdb28ebe63e6fff90ac41f3fcff91
ocr/deu.traineddata.gz1.3 MB306c4280d0cbed46fbff727486bd43b92730181bae80f56941a091f363bdf28b
ocr/spa.traineddata.gz2.1 MB40be52f97b5d4eb7460073dc1f94cd546b27150333c0bf854ed7e7132db6bceb
ocr/ita.traineddata.gz1.7 MBf702fcfad297ce028ede3626d1467b67939f23ff23595f9badd54681cf25a4d3
ocr/nld.traineddata.gz3 MBa2d904b6ddc4feb0d31ecfcd7361a554102e7aa2e278c54f4fc029e0d0815571

The output check

Before offering you the file, the tool reads it back by paths independent of the one that wrote it: the text, with a second PDF reader; the file's objects, walked through one by one, each stream decompressed. Each redacted item is searched for in all its forms (spaces, accents, capitals). The structure is checked: no annotation, no form, no bookmark, no attachment, no metadata, and a single version.

If it finds anything at all, the file is not offered, and the screen says what was found and where.

What is not guaranteed

  • Browser extensions can read the pages you open. For a sensitive document, use a profile with no extensions.
  • A compromised computer (spyware, screen capture) sees everything you see. No website can do anything about it.
  • The browser itself sends information to its publisher (updates, crash reports) depending on its settings.
  • Context: a document whose names are redacted can remain recognisable. See pseudonymisation or anonymisation.
  • Detection suggests; it does not guarantee it has seen everything. The reread is yours.