Skip to content
ConceptThe browser prototype is live — try it on your own text

Unstructured data in.
Validated JSON out.

DataAnchor is a developer platform in the making that turns documents, emails and support tickets into clean, schema-validated JSON — with sensitive data redacted before it reaches your models, APIs or databases.

Built for developers · Designed for control

  • Demo runs in your browser
  • Your text never leaves the tab
  • No sign-up, no tracking
dataanchor / pipeline / invoice
Interactive prototypeAn illustrative view of the planned pipeline. The JSON shown is real output from the same in-browser rules engine that powers the live demo.

Sources

DataAnchor

rules-v1 · in-browser

  1. Detect
  2. Redact
  3. Extract
  4. Validate
– redacted– fields
output.jsoninvoice@v1
processedschema-validated

01The problem

Your data wasn’t built for AI.

Before any model or API can use business information, someone has to find it, clean it, make it safe and force it into a consistent shape. Today that work is repeated, by hand, in every integration.

01

Fragmented and unstructured

Business knowledge is spread across PDFs, inboxes, chat threads and ticket queues — formats written for people, not programs.

02

Sensitive by default

Real messages carry names, emails, phone numbers and payment details. Forwarding them to a model or vendor unfiltered is a liability.

03

Inconsistent structures

The same date or amount arrives in a dozen shapes. Every integration ends up with its own brittle parsing and validation code.

The proposed solution

DataAnchor is the layer in between.

One predictable step between raw inputs and everything that consumes them: detect what is there, remove what should not travel, extract what matters, and validate it against a schema you control.

  • One versioned schema per document type
  • Sensitive values replaced before data moves
  • Every field traceable to the rule that produced it
  • Documents
  • Email
  • Tickets
  • Messages
DataAnchor
  • AI applications
  • APIs & webhooks
  • Databases

02Capabilities

Four steps from raw text to dependable data.

Each capability is labelled with its real status. The working parts run in your browser today; the rest describes where the product is headed.

01 — Extract

Intelligent data extraction

Turn free text into typed fields. Each document type maps to a versioned schema, and every value can be traced back to the text it came from — hover a field to see its source.

Working in prototype

In the prototype

  • Rule-based extraction for invoices, tickets and customer emails
  • Dates, amounts and currencies normalised to ISO formats
  • Fields omitted — never guessed — when no rule matches

Planned

  • Model-assisted extraction for long, irregular documents
  • PDF, scanned-image and attachment parsing
  • Per-field confidence scores and source citations

input.txt

Hi, this is Sarah Mitchell from Northstar Labs. Please process invoice INV-2048 for $1,250, due October 30, 2026. You can contact me at sarah@example.com.

output.json

02 — Redact

Sensitive data redaction

Recognisable identifiers are detected and replaced before data moves on. Placeholders stay consistent within a document, so [NAME_1] refers to the same person everywhere it appears.

Working in prototype

In the prototype

  • Emails, phone numbers, IPv4, US SSNs and common API-key formats
  • Payment cards and IBANs confirmed by checksum
  • Replace, partial-mask or pass-through modes

Planned

  • Statistical entity recognition for names and addresses
  • Per-field and per-destination redaction policies
  • Reversible tokenisation via an access-controlled vault

Prototype name detection relies on context such as “this is …”, “Reported by:” or a sign-off, and will miss names that appear elsewhere. Do not use it as a compliance control.

redaction preview

Reported by: [NAME_1] ([EMAIL_1])
Callback [PHONE_1] · last seen from [IP_1]
Card pasted in chat: [CARD_1]
Refund to IBAN [IBAN_1]
  • Person name

    context rule

  • Email address

    pattern

  • Phone number

    pattern + shape rules

  • IP address

    pattern + octet ranges

  • Payment card

    pattern + Luhn checksum

  • IBAN

    pattern + mod-97 checksum

03 — Validate

Schema mapping and validation

Output conforms to a schema you can read. Required fields, types and enums are checked on every run, and anything that fails is flagged for review instead of silently passing through.

Working in prototype

In the prototype

  • Three built-in schemas plus a generic fallback
  • Type, enum, format and required-field checks
  • processed / needs_review status on every result

Planned

  • Custom schemas defined with JSON Schema
  • Schema versioning and migration
  • Machine-readable validation errors over the API
processed

input

Hi, this is Sarah Mitchell from Northstar Labs. Please process invoice INV-2048 for $1,250, due October 30, 2026. You can contact me at sarah@example.com.

Validation of complete invoice against invoice@v1
fieldtypevaluecheck
document_type*enuminvoiceValid: Classified as invoice
invoice_id*stringINV-2048Valid: Valid string
organizationstringNorthstar LabsValid: Valid string
contact_namestring[REDACTED]Valid, redacted: Valid · redacted in output
contact_emailemail[REDACTED]Valid, redacted: Valid · redacted in output
amount*number1250from “$1,250”Valid: Valid number
currency*currencyUSDfrom “$1,250”Valid: Valid · inferred from “$” — assumed USD
due_datedate2026-10-30from “October 30, 2026”Valid: Valid date
invoice@v1required 4/4 passed

04 — Route

Developer-friendly routing

Validated records should go where they belong. Declarative rules would match on fields and status, so invoices reach finance, tickets reach the help desk and anything uncertain lands in a review queue.

Concept — not implemented

In the prototype

  • Not implemented — the diagram illustrates the intended design

Planned

  • Declarative rules evaluated against validated output
  • Delivery to webhooks, queues and databases
  • Delivery logs, retries and replay
design concept

validated record

{ "document_type": "invoice", "status": "processed" }

routes.yaml · first match wins

  1. 1when status == "needs_review"queue.review
  2. 2when document_type == "invoice"postgres.invoices
  3. 3when document_type == "support_ticket"webhook.helpdesk
  4. 4otherwisearchive
  • Review queue

    human-in-the-loop

  • Postgres(matched)

    finance.invoices

  • Webhook

    POST /helpdesk/tickets

  • Object storage

    archive/

03Live demo

Try the pipeline on your own text.

A working, browser-based version of the DataAnchor pipeline. Load a sample or paste something messy — nothing is uploaded, logged or stored.

dataanchor / playground
prototypeProcessing runs in this tab with deterministic rules. Your text is not sent to any server, not stored, and is gone when you reload.

256 / 20,000 characters. Press Control or Command plus Enter to process.

Redaction mode

Replace identifiers with typed placeholders such as [EMAIL_1].

Pipeline

  1. Analyzing input — pending

  2. Detecting recognizable fields — pending

  3. Redacting sensitive information — pending

  4. Generating structured output — pending

Deterministic rules · no model calls · no network I/O

Structured output will appear here

Load a sample or paste your own text, then press Process data.

Ctrl / ⌘ + Enter runs the pipeline

Deterministic, not generative

The prototype uses regular expressions, checksums and keyword scoring. There is no language model behind it, so identical input always gives identical output.

Local by construction

Everything runs in this tab. The site’s content security policy blocks requests to other origins, and the console reports what it measures.

Honest about gaps

Fields are only emitted when a rule matches. Missing or malformed values are flagged for review instead of being filled in.

Supports English text and three document types. See the engine reference for every rule and its limitations.

04How it works

A pipeline you can reason about.

Six explicit stages, each with a single responsibility and an inspectable output. Hover or select a stage to see what it does and how far along it is.

Detect

Working in prototype

Find every recognisable identifier and candidate field before anything else touches the data.

  • Pattern detectors for emails, phones, IPs and IDs
  • Checksums (Luhn, mod-97) reject look-alike numbers
  • Context rules for names and organisations
  • Rule-based classification with visible signals

Prototype status: Eight detector types and three document classes run in the browser demo.

Detections with their method

"sarah@example.com"     → email    pattern"DE89 3704 … 0130 00"   → iban     mod-97 ✓"4242 4242 4242 4242"   → card     Luhn ✓"Sarah Mitchell"        → person   context: "this is …"

Why this order? Identifiers are removed before extraction so that the planned model-assisted stage never sees raw personal data. In the browser prototype, rules read the original text locally and only redacted values are written to the output.

05Developer experience

Designed to disappear into your stack.

The intended interface is deliberately small: one function, a schema name and a redaction policy. Everything else is data you can inspect.

API in design
  • One call, typed result

    Send text and a schema name. Get back data, redactions and a validation status in one predictable envelope.

  • Explicit failure modes

    needs_review is a first-class outcome with field-level reasons — not an exception you discover in production.

  • Same contract everywhere

    The SDKs, the HTTP API and future webhooks are intended to share one JSON contract.

Illustrative API — planned interface. No SDK, package or endpoint has been published; names and parameters are placeholders and will change. The response values below were generated by the browser prototype engine.

from dataanchor import DataAnchor client = DataAnchor(api_key="YOUR_API_KEY") result = client.extract(    input="Invoice INV-2048 from Northstar Labs: $1,250, "          "due October 30, 2026. Contact: sarah@example.com",    schema="invoice@v1",    redact_pii=True,) if result.status == "needs_review":    send_to_review_queue(result)else:    print(result.data["invoice_id"])  # "INV-2048"    print(result.redactions)          # [Redaction(type="email", …)]

Planned interface · subject to change

06Built for developers

The principles we are building toward.

These are design priorities and product goals that guide every decision — not claims about a production system that does not exist yet.

API-first architecture

Every capability is designed as an API before it gets an interface, so anything a dashboard can do, your code can do too.

Predictable JSON output

Same schema version in, same shape out. Values are typed and normalised, and fields are never silently invented.

Flexible data schemas

Built-in schemas for common documents today; custom, versioned JSON Schemas are the next step.

Privacy-conscious processing

Redaction runs before extraction. Raw identifiers should never need to reach a model, a vendor or a log line.

Modular integrations

Each stage is separable. Use detection on its own, or hand validated output to the queue or database you already run.

Observability and clear errors

Every result should explain itself: which rules fired, what was redacted, and exactly why a record needs review. The demo’s schema-checks view is an early version of this.

07Prototype status

What exists today — and what doesn’t, yet.

DataAnchor is an early-stage concept. Here is an exact account of what has been built so far.

Built and working

  • This website and its documentation
  • In-browser rules engine: analyze → detect → redact → generate
  • Eight identifier detectors, two verified by checksum
  • Three document schemas plus a generic fallback
  • Automated tests covering the engine

In design

  • HTTP API and SDK interface
  • Custom schema format based on JSON Schema
  • Redaction policy model

Planned

  • Model-assisted extraction for PDFs, scans and long documents
  • Routing and delivery to webhooks, queues and databases
  • Hosted processing with audit logs
  • Reversible tokenisation vault

No customers, partners, investors or certifications are claimed, and no delivery dates are published until they can be committed to.

Make your data ready for what comes next.

Explore how structured, privacy-conscious data processing could simplify the next generation of AI applications.