Before you can protect sensitive data, you have to find it. TruePrivacy connects to your cloud storage, databases, and SaaS tools, scans them automatically, and classifies every category of personal data it finds — no engineering work required for most integrations.
Data discovery dashboard showing connected stores and classification results

Connecting data stores

1

Choose a connector

Under DSPM → Data Stores → Connect, pick from 50+ pre-built connectors — SaaS tools such as Salesforce, HubSpot, Slack, Zendesk, and Google Workspace; cloud storage including AWS S3, Google Cloud Storage, and Azure Blob Storage; and databases such as PostgreSQL, MySQL, and MongoDB.
2

Authorize access

Most connectors use OAuth — sign in, grant read access, done. Databases and some cloud services use scoped credentials or API keys instead.
3

Set the scan scope

Include or exclude specific buckets, folders, schemas, or tables. Exclusions are useful for developer sandboxes, anonymized test environments, or systems governed by a separate process.
4

Run the first scan

Trigger the initial discovery scan. It typically completes within 2–4 hours for most organizations; very large environments may take up to 24 hours for the first full pass.
TruePrivacy stores only metadata — where personal data lives, the categories detected, and discovery timestamps. Raw personal data is sampled in place and never copied out of your environment.

Automated discovery scans

Scan typeWhen it runsWhat it does
Initial scanOn connectionFull inventory of the data store and its contents
Incremental scanOn schedule (daily, weekly, or custom)Re-evaluates only changed data, so runs are fast
On-demand scanWhenever you trigger itAd-hoc scan after onboarding a new tool or during an incident
Scheduled rescans detect new data appearing in existing systems and alert your privacy team when personal data shows up in unexpected locations — including shadow data stores created outside standard provisioning.

Classification of PII and sensitive data

Every field, table, and file is evaluated with AI models and configurable rules that detect 100+ types of personal data — names, emails, phone numbers, national IDs, financial identifiers, health data, and biometric references.
  • Data categories — from basic contact details up to GDPR Article 9 special category data and India DPDP sensitive personal data
  • Confidence scores — ambiguous fields (for example, a column named ref_code) are classified from sampled values and surrounding context; low-confidence labels are flagged for human review rather than applied silently
  • Custom taxonomies — add your own categories and detection rules (keywords or patterns) alongside the built-in ones
  • Regulation mapping — each label maps to the regulatory provisions it implicates, led by GDPR, then India DPDP, CCPA, and others
Corrections you make while reviewing findings — confirming, dismissing, or relabeling — teach the system your specific data landscape and improve accuracy over time.

Coverage view

The coverage dashboard shows how much of your estate is discovered and classified:
  • Connected vs. unconnected systems, cross-referenced against your data map
  • Scan freshness per store — when it was last scanned and whether it is on a schedule
  • Classification completeness — the share of fields with confirmed labels versus pending review
Classification results flow into the data map, RoPA, and DSR Automation, so every downstream tool works from the same authoritative view of where data lives.
Unscanned stores are invisible risk. If a system holds personal data but is not connected, add it to the data map manually so it is at least inventoried while you work on connectivity.
Once discovery is running, findings about how well that data is protected appear in Risk Findings.