Unstructured data in.
Validated JSON out.
DataAnchor is a developer platform in the making that turns documents, emails and support tickets into clean, schema-validated JSON — with sensitive data redacted before it reaches your models, APIs or databases.
Built for developers · Designed for control
- Demo runs in your browser
- Your text never leaves the tab
- No sign-up, no tracking
Sources
DataAnchor
rules-v1 · in-browser
- Detectpatterns + checksums
- Redacttyped placeholders
- Extractschema fields
- Validatetypes + required
01The problem
Your data wasn’t built for AI.
Before any model or API can use business information, someone has to find it, clean it, make it safe and force it into a consistent shape. Today that work is repeated, by hand, in every integration.
01
Fragmented and unstructured
Business knowledge is spread across PDFs, inboxes, chat threads and ticket queues — formats written for people, not programs.
02
Sensitive by default
Real messages carry names, emails, phone numbers and payment details. Forwarding them to a model or vendor unfiltered is a liability.
03
Inconsistent structures
The same date or amount arrives in a dozen shapes. Every integration ends up with its own brittle parsing and validation code.
The proposed solution
DataAnchor is the layer in between.
One predictable step between raw inputs and everything that consumes them: detect what is there, remove what should not travel, extract what matters, and validate it against a schema you control.
- One versioned schema per document type
- Sensitive values replaced before data moves
- Every field traceable to the rule that produced it
- Documents
- Tickets
- Messages
extract · validate
- AI applications
- APIs & webhooks
- Databases
02Capabilities
Four steps from raw text to dependable data.
Each capability is labelled with its real status. The working parts run in your browser today; the rest describes where the product is headed.
01 — Extract
Intelligent data extraction
Turn free text into typed fields. Each document type maps to a versioned schema, and every value can be traced back to the text it came from — hover a field to see its source.
Working in prototypeIn the prototype
- Rule-based extraction for invoices, tickets and customer emails
- Dates, amounts and currencies normalised to ISO formats
- Fields omitted — never guessed — when no rule matches
Planned
- Model-assisted extraction for long, irregular documents
- PDF, scanned-image and attachment parsing
- Per-field confidence scores and source citations
input.txt
Hi, this is Sarah Mitchell from Northstar Labs. Please process invoice INV-2048 for $1,250, due October 30, 2026. You can contact me at sarah@example.com.
output.json
02 — Redact
Sensitive data redaction
Recognisable identifiers are detected and replaced before data moves on. Placeholders stay consistent within a document, so [NAME_1] refers to the same person everywhere it appears.
Working in prototypeIn the prototype
- Emails, phone numbers, IPv4, US SSNs and common API-key formats
- Payment cards and IBANs confirmed by checksum
- Replace, partial-mask or pass-through modes
Planned
- Statistical entity recognition for names and addresses
- Per-field and per-destination redaction policies
- Reversible tokenisation via an access-controlled vault
Prototype name detection relies on context such as “this is …”, “Reported by:” or a sign-off, and will miss names that appear elsewhere. Do not use it as a compliance control.
redaction preview
Reported by: [NAME_1] ([EMAIL_1]) Callback [PHONE_1] · last seen from [IP_1] Card pasted in chat: [CARD_1] Refund to IBAN [IBAN_1]
Person name
context rule
Email address
pattern
Phone number
pattern + shape rules
IP address
pattern + octet ranges
Payment card
pattern + Luhn checksum
IBAN
pattern + mod-97 checksum
03 — Validate
Schema mapping and validation
Output conforms to a schema you can read. Required fields, types and enums are checked on every run, and anything that fails is flagged for review instead of silently passing through.
Working in prototypeIn the prototype
- Three built-in schemas plus a generic fallback
- Type, enum, format and required-field checks
- processed / needs_review status on every result
Planned
- Custom schemas defined with JSON Schema
- Schema versioning and migration
- Machine-readable validation errors over the API
input
Hi, this is Sarah Mitchell from Northstar Labs. Please process invoice INV-2048 for $1,250, due October 30, 2026. You can contact me at sarah@example.com.
04 — Route
Developer-friendly routing
Validated records should go where they belong. Declarative rules would match on fields and status, so invoices reach finance, tickets reach the help desk and anything uncertain lands in a review queue.
Concept — not implementedIn the prototype
- Not implemented — the diagram illustrates the intended design
Planned
- Declarative rules evaluated against validated output
- Delivery to webhooks, queues and databases
- Delivery logs, retries and replay
validated record
{ "document_type": "invoice", "status": "processed" }
routes.yaml · first match wins
- 1when status == "needs_review"queue.review
- 2when document_type == "invoice"postgres.invoices
- 3when document_type == "support_ticket"webhook.helpdesk
- 4otherwisearchive
Review queue
human-in-the-loop
Postgres(matched)
finance.invoices
Webhook
POST /helpdesk/tickets
Object storage
archive/
03Live demo
Try the pipeline on your own text.
A working, browser-based version of the DataAnchor pipeline. Load a sample or paste something messy — nothing is uploaded, logged or stored.
256 / 20,000 characters. Press Control or Command plus Enter to process.
Redaction mode
Replace identifiers with typed placeholders such as [EMAIL_1].
Pipeline
Analyzing input — pending
Detecting recognizable fields — pending
Redacting sensitive information — pending
Generating structured output — pending
Deterministic rules · no model calls · no network I/O
Structured output will appear here
Load a sample or paste your own text, then press Process data.
Ctrl / ⌘ + Enter runs the pipeline
Deterministic, not generative
The prototype uses regular expressions, checksums and keyword scoring. There is no language model behind it, so identical input always gives identical output.
Local by construction
Everything runs in this tab. The site’s content security policy blocks requests to other origins, and the console reports what it measures.
Honest about gaps
Fields are only emitted when a rule matches. Missing or malformed values are flagged for review instead of being filled in.
Supports English text and three document types. See the engine reference for every rule and its limitations.
04How it works
A pipeline you can reason about.
Six explicit stages, each with a single responsibility and an inspectable output. Hover or select a stage to see what it does and how far along it is.
Detect
Working in prototypeFind every recognisable identifier and candidate field before anything else touches the data.
- Pattern detectors for emails, phones, IPs and IDs
- Checksums (Luhn, mod-97) reject look-alike numbers
- Context rules for names and organisations
- Rule-based classification with visible signals
Prototype status: Eight detector types and three document classes run in the browser demo.
Detections with their method
"sarah@example.com" → email pattern"DE89 3704 … 0130 00" → iban mod-97 ✓"4242 4242 4242 4242" → card Luhn ✓"Sarah Mitchell" → person context: "this is …"Why this order? Identifiers are removed before extraction so that the planned model-assisted stage never sees raw personal data. In the browser prototype, rules read the original text locally and only redacted values are written to the output.
05Developer experience
Designed to disappear into your stack.
The intended interface is deliberately small: one function, a schema name and a redaction policy. Everything else is data you can inspect.
One call, typed result
Send text and a schema name. Get back data, redactions and a validation status in one predictable envelope.
Explicit failure modes
needs_review is a first-class outcome with field-level reasons — not an exception you discover in production.
Same contract everywhere
The SDKs, the HTTP API and future webhooks are intended to share one JSON contract.
Illustrative API — planned interface. No SDK, package or endpoint has been published; names and parameters are placeholders and will change. The response values below were generated by the browser prototype engine.
from dataanchor import DataAnchor client = DataAnchor(api_key="YOUR_API_KEY") result = client.extract( input="Invoice INV-2048 from Northstar Labs: $1,250, " "due October 30, 2026. Contact: sarah@example.com", schema="invoice@v1", redact_pii=True,) if result.status == "needs_review": send_to_review_queue(result)else: print(result.data["invoice_id"]) # "INV-2048" print(result.redactions) # [Redaction(type="email", …)]Planned interface · subject to change
06Built for developers
The principles we are building toward.
These are design priorities and product goals that guide every decision — not claims about a production system that does not exist yet.
API-first architecture
Every capability is designed as an API before it gets an interface, so anything a dashboard can do, your code can do too.
Predictable JSON output
Same schema version in, same shape out. Values are typed and normalised, and fields are never silently invented.
Flexible data schemas
Built-in schemas for common documents today; custom, versioned JSON Schemas are the next step.
Privacy-conscious processing
Redaction runs before extraction. Raw identifiers should never need to reach a model, a vendor or a log line.
Modular integrations
Each stage is separable. Use detection on its own, or hand validated output to the queue or database you already run.
Observability and clear errors
Every result should explain itself: which rules fired, what was redacted, and exactly why a record needs review. The demo’s schema-checks view is an early version of this.
07Prototype status
What exists today — and what doesn’t, yet.
DataAnchor is an early-stage concept. Here is an exact account of what has been built so far.
Built and working
- This website and its documentation
- In-browser rules engine: analyze → detect → redact → generate
- Eight identifier detectors, two verified by checksum
- Three document schemas plus a generic fallback
- Automated tests covering the engine
In design
- HTTP API and SDK interface
- Custom schema format based on JSON Schema
- Redaction policy model
Planned
- Model-assisted extraction for PDFs, scans and long documents
- Routing and delivery to webhooks, queues and databases
- Hosted processing with audit logs
- Reversible tokenisation vault
No customers, partners, investors or certifications are claimed, and no delivery dates are published until they can be committed to.
Make your data ready for what comes next.
Explore how structured, privacy-conscious data processing could simplify the next generation of AI applications.