Almost every business database holds personal data: customer names, phone numbers, email and postal addresses, order histories, sometimes identity numbers or health details. Protecting it is a legal duty in many countries and a matter of customer trust everywhere. This article focuses on the practical, technical side of personal data protection inside your database: what to do with tables, columns, access and copies. It is general good practice, not legal advice. Laws differ between countries, and you should check your specific obligations with a qualified adviser.
Personal data protection: a note on the legal landscape
Many jurisdictions now have data protection laws. The European Union's General Data Protection Regulation (GDPR) applies to organisations handling personal data of people in the EU, including many businesses based elsewhere. India's Digital Personal Data Protection Act, 2023 (DPDP Act) sets out obligations for organisations processing digital personal data in India. Other countries and some US states have their own rules.
The details, such as what counts as personal data, consent requirements, breach notification and penalties, vary. But several principles appear in most frameworks: collect only what you need, use it only for stated purposes, keep it secure, do not keep it longer than necessary, and be able to respond when people ask about their data. Every step below supports one of those principles.
Step 1: Know where personal data lives
You cannot protect data you do not know you hold. Build a simple data inventory: a list of tables and columns containing personal data, what category each is, and why you hold it. Start by searching the schema for likely column names:
SELECT table_name, column_name, data_type
FROM information_schema.columns
WHERE table_schema = 'shopdb'
AND column_name REGEXP 'name|email|phone|mobile|address|dob|birth|pan|aadhaar|passport|ip_addr'
ORDER BY table_name;
(This uses MySQL's REGEXP; in PostgreSQL use ~*.) Then look further than the obvious: free-text notes fields where staff paste details, JSON columns, log tables that capture IP addresses, file uploads, and email archives. Do not forget copies: backups, analytics databases, spreadsheets exported by staff and test environments.
A useful inventory format:
| Location | Data | Purpose | Sensitivity | Retention |
|---|---|---|---|---|
customers.email | Email address | Order updates, login | Normal | While account is active |
kyc_documents.id_number | Government ID number | Identity verification | High | As required by applicable rules |
access_logs.ip_address | IP address | Security monitoring | Normal | Short, for example 90 days |
Step 2: Collect and keep less
Data minimisation means holding only what you genuinely need. Every field you do not store is a field that cannot leak. Ask of each column: do we use this? If you collect date of birth only to confirm a customer is an adult, perhaps a confirmed flag is enough. If you store full card numbers, stop: payment gateways can tokenise cards so you never handle them, and card data falls under separate industry security standards.
Step 3: Restrict who can see what
Not everyone who uses the database needs to see personal details. Practical techniques include:
- Separate database accounts for the application, reporting and administrators, each with only the privileges it needs.
- Views that expose only safe columns to analysts:
CREATE VIEW customers_for_analytics AS
SELECT id,
city,
state_code,
YEAR(created_at) AS signup_year,
SHA2(CONCAT(email, 'per-project-secret'), 256) AS email_hash
FROM customers;
GRANT SELECT ON shopdb.customers_for_analytics TO 'analyst'@'10.0.2.%';
- Row-level security (built into PostgreSQL) so that, for example, regional staff only see their region's customers.
- Application-level roles that hide sensitive fields from screens where they are not needed.
The hashed email in the view above is an example of pseudonymisation: replacing a direct identifier with a value that still lets analysts count and join records without seeing the email itself. Pseudonymised data is generally still treated as personal data under laws such as the GDPR, because it can be linked back, but it does reduce exposure.
Step 4: Encrypt, especially the sensitive fields
Encrypt connections with TLS and encrypt storage at rest. For high-sensitivity values such as identity numbers, consider application-level encryption: the application encrypts the value before saving it, and the keys live in a key management service outside the database. Then a stolen database dump, on its own, does not reveal those values. Store passwords only as slow, salted hashes, never in reversible form.
Step 5: Never use raw production data for testing
Development and test databases are often less protected than production and accessed by more people, including contractors. Copying production into them spreads personal data widely. Use data masking instead: replace real values with realistic fakes while keeping formats and relationships intact.
UPDATE customers
SET name = CONCAT('Customer ', id),
email = CONCAT('user', id, '@example.test'),
phone = CONCAT('90000', LPAD(id % 100000, 5, '0'));
Run masking as part of the process that creates the test copy, so unmasked data never lands in the test environment at all.
Step 6: Delete on schedule, and on request
Define how long each kind of data is kept, then enforce it with scheduled jobs rather than good intentions. Some records, such as invoices, may need to be kept for a period set by tax or accounting rules, while the marketing profile linked to them can be removed sooner. Plan how you will:
- Delete or anonymise a customer when retention ends or they ask, including related tables.
- Export a person's data if they request a copy.
- Correct inaccurate data.
- Handle backups, which usually cannot be edited; many organisations let deleted data age out of backups on the normal rotation and make sure it is not restored back into live systems.
Anonymisation must be genuine: removing the name but keeping the exact address and date of birth may still identify someone.
Step 7: Log access and prepare for incidents
Record who accesses or exports personal data, and alert on unusual volumes. Write down, before you need it, what you will do if a breach happens: who investigates, who decides on notification, and where the inventory from Step 1 lives. Notification deadlines under some laws are short. Our database management team helps implement masking, access controls and retention jobs, and you can generate strong credentials for database accounts with our password generator.
Key takeaways
- Personal data protection starts with an inventory of where personal data lives, including copies.
- Collect less, restrict access with accounts, views and row-level security, and encrypt sensitive fields.
- Mask data before it reaches test environments.
- Automate retention and deletion, and check your obligations under the laws that apply to you.