Privacy by default: how we think about data in AI products
Privacy isn't a settings screen. It's architecture decisions made before writing the first line of code.
- Category
- Engineering
- Author
- Equipe Devx
- Reading time
- 1 min
It's common to treat privacy as a layer added at the end: a policy written after the product is done, a delete-account button, a cookie banner. In products that process sensitive information, that doesn't work. The decisions that matter most are made much earlier.
Four questions before the code
Every feature that touches user data goes through four questions:
- Do we really need this data? The safest data is the data that was never collected.
- For how long? Retention has a defined period, and the default period is the shortest possible.
- Where is it processed? Knowing exactly which services touch the data — and being able to explain that to anyone who asks.
- Who can see it? Minimum access by default, including for our own team.
A concrete example: this website
The Devx website itself follows these rules. Audience measurement uses no cookies and doesn't store IP addresses. Each visit generates an identifier derived from technical request data combined with a secret that changes every day — enough to count unique visitors in a day, not enough to follow anyone over time.
The result is that we know how many people visited a page and how many clicked on a product, but we don't know who they are. For the decisions we need to make, that's exactly enough.
And in AI products
Language models create a new temptation: send everything to an external service because it's easier. In our products, the question "where is this processed?" applies to models too. When using an external service makes sense, that will be documented — and when it doesn't, we build the alternative.
- #privacy
- #lgpd
- #architecture