| title | Migrate from DataFog Python |
|---|---|
| description | Move from datafog 4.8.x to the canonical DataFog Core 0.3.x Python API. |
| icon | arrow-right-arrow-left |
DataFog Core makes detection results, text ranges, transformation policy, and provider-backed operations consistent across Rust, Python, Node.js, and browser/WASM. That consistency requires a few deliberate API changes in Python.
DataFog Core 0.3 does not replace the legacy package's optional spaCy, GLiNER, OCR, distributed-processing, CLI, or application-guardrail features. Keep the established package for those workloads while adopting Core where its smaller, cross-runtime contract fits.
python -m pip install datafogpython -m pip install datafog-corefrom datafog.engine import Entity, redact, scan, scan_and_redactfrom datafog_core import Finding, scan, scan_and_transform, transformThe distribution name uses a hyphen (datafog-core), while the Python import
uses an underscore (datafog_core). The two distributions can be installed at
the same time while an application is migrated incrementally.
DataFog Python returns a ScanResult wrapper. DataFog Core returns a list of
Finding objects directly.
from datafog.engine import scan
result = scan(
"Email jane@example.com",
engine="regex",
entity_types=["EMAIL"],
)
for entity in result.entities:
print(entity.type, entity.text, entity.start, entity.end)from datafog_core import scan
findings = scan("Email jane@example.com")
for finding in findings:
print(
finding.entity_type,
finding.matched_text,
finding.codepoint_range.start,
finding.codepoint_range.end,
)Use codepoint_range when slicing a Python string. Use byte_range when
addressing the UTF-8 encoded input. Both ranges are zero-based and
end-exclusive.
| DataFog Python 4.8.x | DataFog Core 0.3.x |
|---|---|
ScanResult.entities |
The return value from scan |
Entity.type |
Finding.entity_type |
Entity.text |
Finding.matched_text |
Entity.start, Entity.end |
Finding.codepoint_range.start, .end for Python string indexing |
| No distinct UTF-8 range | Finding.byte_range |
Entity.engine |
Finding.detector_name and .detector_version |
ScanResult.engine_used |
No aggregate equivalent; provenance is attached to each finding |
Rule-based Core findings can have confidence=None. Do not assume every
finding has a numeric confidence score.
DataFog Core separates detection configuration from transformation policy.
transform requires explicit findings; scan_and_transform is the convenience
operation that performs both steps.
from datafog.engine import scan_and_redact
result = scan_and_redact(
"Email jane@example.com",
engine="regex",
entity_types=["EMAIL"],
strategy="mask",
)
print(result.redacted_text)from datafog_core import scan_and_transform
result = scan_and_transform(
"Email jane@example.com",
{
"transform": {
"default": {"strategy": "mask"},
"entities": ["EMAIL"],
}
},
)
print(result.text)| DataFog Python 4.8.x | DataFog Core 0.3.x |
|---|---|
scan_and_redact(...) |
scan_and_transform(text, {"scan": ..., "transform": ...}) |
redact(text, entities, ...) |
transform(text, findings, config) |
RedactResult.redacted_text |
TransformResult.text |
RedactResult.entities |
TransformResult.transformations |
RedactResult.mapping |
No plaintext mapping is returned |
Transformation records describe what was applied, including source and output ranges, but intentionally omit the original matched PII.
Do not migrate strategy names mechanically. In particular, legacy token and
Core tokenize have different security and reversibility semantics.
| DataFog Python 4.8.x strategy | DataFog Core migration |
|---|---|
mask |
Use mask. Core also supports explicit leading or trailing reveal rules. |
token |
For an irreversible type placeholder, use redact. Use tokenize only when provider-backed restoration is required. |
hash |
There is no unkeyed hash strategy. Use keyed pseudonymize only when stable, non-reversible linkage is required. |
pseudonymize |
Review the behavior rather than renaming it. Core pseudonymization is deterministic, keyed, and provider-backed. |
Core also adds remove, which deletes only the exact finding span.
See Privacy transformations for the full behavior and threat-model distinctions.
Do not translate engine="regex", "smart", "spacy", or "gliner" into
Core configuration. DataFog Core 0.3 owns detector composition and exposes
locale as its scan setting. Detector provenance appears on each finding.
Entity names are exact and case-sensitive. The built-in Core entities are:
EMAIL, PHONE, SSN, CREDIT_CARD, IP_ADDRESS, DATE, and ZIP_CODE.
Use canonical names such as DATE and ZIP_CODE rather than legacy aliases
such as DOB or ZIP. Entity selection belongs in the transformation config,
not the scan config.
- Replace the
datafogdistribution anddatafog.engineimports. - Consume the list returned by
scaninstead ofScanResult.entities. - Rename entity fields and select
codepoint_rangeorbyte_rangeexplicitly. - Replace
scan_and_redactwithscan_and_transformand a transformation envelope. - Remove legacy engine selectors and normalize entity names.
- Select a Core strategy by behavior, especially for
tokenandpseudonymize. - Stop depending on plaintext mappings or original PII in transformation records.
- Update error handling for the DataFog Core exceptions.
Continue with the Python reference, then review findings and ranges and configuration.
The seven German labels are implemented in Core, with shared Rust, installed
Python/Node and browser/WASM conformance fixtures. This is format/context
compatibility, not exact regex parity with Python 4.8.1. The comparison baseline
is the published 4.8.1 contract and its tests/test_de_pii_regex.py; those legacy
fixtures remain unchanged. fixtures/german.jsonl records the new contract.
| Behavior | Python 4.8.1 comparison | Core policy |
|---|---|---|
| Activation | German locale or explicit German entity selection | Singular locale; transformation selection alone does not activate a detector |
| IBAN validation | Format detection, no checksum requirement | Same; checksum-invalid candidates are detected |
| Whitespace | Broad regex \s acceptance |
Only space, tab, NBSP, narrow NBSP at prescribed positions; newline and other Unicode whitespace rejected |
| Digits | Unicode regex \d |
ASCII [0-9] only |
| Context | Legacy context-substring matching | Immediate context + horizontal gap + value; ASCII boundary on context and value; no intervening prose/newline |
| Grouping | Legacy regex forms | Independent optional single separators at documented boundaries; no arbitrary or repeated value separators |
| Postal spans | Prefix included | Prefix included; PLZ: 10115 explicitly rejected |
| Passport/residence | One-letter/eight-digit and AT/seven-digit heuristics | Same limited heuristics, not exhaustive official document coverage |
| Generic overlaps | Adapter selection behavior is legacy-specific | Scans retain both candidates; transformations use existing Core overlap rules |
All matching preserves source casing and separators. Allowlist entries must use that source representation. Required context stays in the output, except postal prefixes, which are part of the finding and are replaced too.
The Python adapter may translate explicit German entity requests into locale
de and then filter results, but must separately test legacy overlap and
selection semantics. Core adds no plural locales, engine switch or scan-time
entity selector. The capability-aware higher-level adapter and its
datafog-core>=0.4.0,<0.5 dependency were merged in
PR #179. The higher-level
Python package has its own release process; see the
reviewed legacy NPI selection limitation.
Core 0.4.1 adds API keys, explicit Bearer header tokens, and scoped PostgreSQL credential URIs. Python 4.9 retains its existing default backend and legacy overlap-before-selection behavior. A same-span API key can suppress a Bearer token, a Bearer token can suppress a JWT, and a containing credential URI can suppress an inner API key or JWT before legacy selection runs. Explicit selection can therefore return empty. Native Core and datafog.v5 retain scan candidates and apply selection before overlap during transformation. See the compatibility matrix and exact-wheel integration results.