Privacy as a Common Good: Data Rights in the AI Era
GDPR, Japan's APPI, US state laws and China's PIPL in 2026, plus the five privacy-enhancing technologies now in commercial use and the self-sovereign identity systems being built on them.
How Much Data, and Who Collects It
An average internet user generates about 1.7 megabytes of data every second. Across five billion connected people that comes to more than 400 exabytes a day. Most of it is personal or behavioral, and most of it is monetized under consent terms few people have read.
Large language models and multimodal AI have raised the demand. A frontier model needs variety as well as volume: conversation, cultural context, professional knowledge. Chatbot queries, documents uploaded to cloud services and biometric scans at airport checkpoints all end up in training sets.
Advertising based on behavioral prediction brought in more than $600 billion worldwide in 2025.
More than 8.2 billion records were exposed in publicly disclosed breaches in 2025. The figure covers only incidents that organizations were required or chose to report. The average cost per breach has reached $4.9 million, a number that does not include lost public trust.
Regulation: Similar Principles, Different Practice
Europe’s GDPR has been in force for eight years. Cumulative fines passed EUR 4.5 billion by the end of 2025, with large penalties against Meta, Amazon and TikTok. The EU’s AI Act began phased enforcement in 2025 and adds transparency and data governance requirements for high-risk AI systems. The EU Digital Identity Wallet, expected in broad deployment by 2027, would let citizens present verified credentials without handing over the personal information behind them.
Japan’s APPI (Act on the Protection of Personal Information) was amended substantially in 2024-2025: tighter rules on cross-border transfers, pseudonymized data and the right to deletion. Japan holds an adequacy agreement with the EU, one of only a handful worldwide, so Japanese companies can operate under both regimes without extra friction. The Digital Agency, set up in 2021 under then-Minister Taro Kono, has been modernizing digital identity infrastructure, including wider use of the My Number system and pilots for verifiable digital credentials.
The United States has no federal privacy law. The rules come from states. California has CCPA/CPRA, Colorado the CPA, and Connecticut, Virginia and a growing list of others have their own, each with different thresholds, definitions and enforcement. A multinational can face up to fifty privacy regimes in one market. The FTC brings enforcement actions under its existing authority. Among advanced economies the U.S. is the one without a single framework.
China’s PIPL (Personal Information Protection Law) has been in force since 2021 and contains one of the strictest consent regimes anywhere. Enforcement has been selective. State-affiliated entities are largely exempt from the constraints on private companies, so citizen data flows to government surveillance while private-sector use is tightly policed.
Privacy-Enhancing Technologies
Five privacy-enhancing technologies (PETs) are in commercial use.
Federated learning trains AI models across separate datasets, on hospital servers, phones or corporate networks, without moving the data. Google uses it at scale for keyboard prediction and health research. Apple builds on-device processing into its architecture. Model updates can still leak information, and coordinating across mismatched systems is hard.
Differential privacy adds calibrated noise to datasets or query results, with a mathematical guarantee that no individual can be reverse-engineered from the aggregate. The U.S. Census Bureau used it for the 2020 Census. Apple and Google apply it to usage analytics. An enterprise that has to share data with partners, regulators or researchers can do so with a defined bound on what leaks.
Homomorphic encryption runs computations on encrypted data without decrypting it. For years it was too slow to use. IBM, Microsoft and Intel have all released libraries with much better performance, and financial institutions are piloting encrypted analytics for fraud detection and credit scoring on data they never see in plaintext.
Zero-knowledge proofs let one party prove a statement (“I am over 18,” “I hold a valid credential,” “my balance exceeds the threshold”) without revealing anything else. They began as a cryptographic curiosity and now sit under working identity and compliance systems. Self-sovereign identity depends on them: a person carries verifiable credentials and discloses only what a given transaction needs.
Synthetic data is generated to match the statistical properties of real data while containing no actual personal records. It is used for AI training, software testing and research. Gartner projects that by 2030 synthetic data will be used more often than real data to train AI models.
Self-Sovereign Identity
In a self-sovereign identity (SSI) system the credential sits in the holder’s own wallet. Nothing leaves it unless the holder agrees.
The EU Digital Identity Wallet is the biggest implementation under way. Each EU citizen would get a government-backed digital identity that works across borders and services, for opening a bank account or proving a professional qualification, with no central database behind it.
Japan’s Digital Agency is piloting digital credentials in healthcare and education on top of the expanded My Number system.
Web3 supplies the plumbing: decentralized identifiers (DIDs), verifiable credentials and blockchain-based attestation. A holder shows a credential to a service, and the service does not keep a copy of the personal data underneath.
Socious Verify is one such system in production. It confirms a qualification, a certification or a record of contribution as a verifiable credential, so the checking party gets a yes or no and not the underlying personal file.
AI Training and Privacy Law
GDPR and APPI both require data minimization and purpose limitation, and both give individuals a right to deletion. Frontier models are trained on the largest and most varied datasets their developers can assemble, and developers keep training data for future versions rather than deleting it. Individual-level behavioral data improves a model more than any other kind. It is also the kind regulators most want pseudonymized or anonymized.
Federated learning and differential privacy take more engineering time than collecting everything into one database. Most organizations still collect everything. In a 2025 Cisco survey, 87% of consumers said they would not do business with a company they did not trust to handle their data responsibly.
Join the Conversation
The Tech for Impact Summit meets on April 26, 2026, at Tokyo Garden Terrace Kioi Conference. This year’s theme is “Beyond Boundaries: Building 2050 Together”. Privacy and digital rights are on the agenda.
Among the confirmed speakers: Taro Kono (former Minister of Digital Affairs), Charles Hoskinson (Cardano founder), Yoshito Hori (GLOBIS), Kathy Matsui (MPower Partners), Ken Suzuki (SmartNews), Sota Watanabe (Astar/Startale), and Jesper Koll (Monex Group).
Executives responsible for AI governance or digital identity at their companies are among those invited.
Explore partnership and membership opportunities →
Watch highlights from previous summits: youtu.be/ujy7ZXflrt4
The Tech for Impact Summit is an invitation-only executive gathering taking place April 26, 2026, in Tokyo as a partner event of SusHi Tech Tokyo. Learn more at tech4impactsummit.com.