PII-Fi Scan

High-speed, high-accuracy detection and masking of Japanese PII and sensitive data

Find 84 types of PII and sensitive information hidden in Japanese text, review them, and hand over pseudonymized files

PII-Fi Scan is a web app that automatically detects 84 types of PII and sensitive information in Japanese text, from names and contact details to API keys and special-category data, and lets you mask or pseudonymize them while reviewing the results. Use it when sharing incident logs, inquiries, meeting minutes or source code outside your organization, or to sanitize text before passing it to an LLM. Everything runs in the browser. No API integration or scripting is needed.

Japanese personal data is hard for general-purpose multilingual tools to catch. Honorifics and job titles become part of a name ("Yamada-bucho", "Tanaka-sama"), whether "Kawasaki" is a person, a place or a company depends on context, and addresses and phone numbers have many notational variants. PII-Fi is built on a detection engine designed specifically for Japanese to tackle these difficulties head-on.

It takes three steps. Drop a set of files (multiple files, folders and ZIP archives are accepted, and the character encoding is detected automatically), review the color-coded detections on screen and correct any hits or misses, then pick a profile and download the pseudonymized files together with an audit report.

Key Features of PII-Fi Scan

Five detection routes working together

Fixed-format items such as phone numbers, email addresses, My Number and API keys are caught by patterns; real personal names, place names, organizations and disease names by dictionaries of about one million entries; surname and given-name combinations by morphological analysis; information that becomes sensitive only in context, such as symptoms or annual income, by context analysis; and names, abbreviations and foreign names absent from dictionaries by a machine-learning NER model trained on Japanese. Conflicting results are resolved automatically by priority and confidence.

A catalog of 84 entity types, configurable per type

40 identifier types whose values point directly to a person or a secret, 35 attribute types that become sensitive in context, and 9 credential types that are dangerous the moment they leak. Credentials are covered by 196 secret-detection rules for services such as AWS, GitHub and Stripe, including dedicated rules for about 35 Japanese SaaS providers. For each type you decide whether to detect it and how to pseudonymize it.

Five pseudonymization methods with consistent mapping

Choose per type from realistic dummy values, type labels with sequence numbers, generalization such as "42 years old" to "40s", redaction, or detect-only without replacement. The same original value maps to the same dummy wherever it appears in a document, so timelines and correlations survive pseudonymization. Phone numbers stay in phone-number format and dates remain valid dates.

Profiles and custom recognizers

Combinations of detected types and methods can be saved as profiles. Four presets ship as standard: log and incident sanitization, source code audit, email sanitization, and minutes for external sharing. Your own management numbers, product code names and internal project names can be added through dictionaries, detection patterns, context-aware custom detectors and exclusion lists.

Designed not to keep your original text

  • Working data is encrypted with a per-job key and stored temporarily. Normal work is deleted 60 minutes after the last operation; explicitly saved work is deleted after 48 hours
  • No restoration material is produced or persisted (there is no function to recover original values from pseudonymized files)
  • Server access logs and job logs record only metrics such as character counts, never the text itself
  • The download ZIP includes an audit report of the processing (a PDF version can be downloaded separately)
  • For customers who cannot send confidential data outside, an on-premises configuration that stays within the internal network is proposed individually
  • Supported files: text-based files such as logs, CSV, JSON, Markdown and source code (10MB per file, up to 200 files per batch, UTF-8 / CP932 / EUC-JP detected automatically)

Pricing (unit-based, counting only the characters processed)

One unit is 100,000 characters, including both detection and pseudonymization. Usage is counted by cumulative characters within the month, so re-scanning small files within the same 100,000-character allowance costs nothing extra. Every Personal plan includes detection, pseudonymization and audit reports for all 84 types. The first contract includes a 14-day free trial after card registration.

PlanMonthlyAnnualMonthly unitsProfilesDetection tuning
Personal Lite19,800 yen198,000 yen100 (about 10 million chars)3Not included
Personal Standard49,800 yen498,000 yen500 (about 50 million chars)10Custom dictionaries, exclusion lists, detection patterns
Personal Pro98,000 yen980,000 yen2,000 (about 200 million chars)50Standard features plus context-aware custom detectors

Monthly units reset every month; shortfalls can be covered with additional unit packs (pricing on request). Team use and on-premises deployment are arranged individually. The service is currently provided through sales consultation. See the plans on the PII-Fi official site (Japanese) for the latest terms (transcribed as of 2026-09-08).

FAQ (excerpt)

Is the uploaded original text stored?

Working data is encrypted with a per-job key and stored temporarily. Normal work is deleted 60 minutes after the last operation, and explicitly saved work after 48 hours. Original text is never stored permanently.

How accurate is detection, and what about false positives?

Please verify with data from your own domain. The first contract includes a 14-day free trial after card registration, and a detection demo with your real data can be arranged during sales consultation. False positives can be tuned with custom dictionaries, exclusion lists and detection patterns (Personal Standard and above) or context-aware custom detectors (Personal Pro), and every hit can be accepted or rejected individually on screen.

What happened to the previous PII-Fi API?

New provision of the former PII-Fi API (v1) has ended, and a new version (v2) is being prepared. If you wish to use PII-Fi through an API, please contact us.

Product Site

The full catalog of 84 types, how detection works and plan details are on the official site (Japanese).

Go to PII-Fi Official Site

Contact Us

Tell us about your data and requirements. A detection demo with your actual files is available

Go to Contact Form