AI-POWERED SENSITIVE INFORMATION TYPE | MICROSOFT PURVIEW
Designing trust into AI-powered data classification
AI-powered sensitive information types help users get more accurate classification, which eventually improves all downstream processes. This case study is about how to help users trust and adopt AI in their work.
MVP · Work is ongoing
WHAT THE DESIGN THINKING ENABLES
01
Understand
Gives admins clear reasons and context behind AI validations so they can make informed decisions.
02
Validate
Brings human-in-the-loop where admins can improve AI performance.
03
Trust
Transparency for users to make investments in AI-powered classification.
01/ CONTEXT
Before we dive in
Read this already? Skip to 02 /problemWhat's Microsoft Purview?
Microsoft Purview is Microsoft's compliance suite for enterprises, a set of tools that help organisations protect sensitive data and meet regulations like HIPAA, GDPR, and others.
What's Data Classification?
Imagine you're moving houses and you're packing your things in boxes.
- Some boxes contain expensive plates, others important files, etc.
- After closing the boxes you don't know what's inside the boxes.
- To help you handle and transport these boxes carefully you'd label certain boxes as fragile, important, books, donate etc.
- Labels help you give appropriate attention to the box.
Purview's Data Classification does the same for organisational data. It scans files and emails, 'labelling' them based on sensitive content like credit card numbers or employee IDs. This labelling tells other compliance tools how to treat each file.
Data classification is a foundational step many other Purview tools rely on.
Items you need to relocate
Packed within boxes
Boxes with appropriate labels
DATA CLASSIFICATION
Scan data in the environment
Apply sensitivity marker
Other Purview tools apply protections to data marked sensitive.
02/ THE PROBLEM
False positives are a recurring pain-point for compliance admins
Compliance admins struggled to distinguish meaningful classifier results from irrelevant or incorrect classifications, creating unnecessary noise in their workflows.
False positives
Data classification system often marks files that are not sensitive as being sensitive
Generates noise
These files show up in the data explorers, flooding their screen.
Manual verification
Human validation is needed to mark the files as false positives.
Lost productivity
Overall takes up time and effort across solutions
False positives mainly occur because the context of the file could not be understood.
For example a file may contain mock credit card numbers for testing, Purview data classification system tags it as sensitive but business requirement says it is not.
THE IDEA
AI-powered validation of the context of a file to weed out false positives at the root during classification.
How might we build trust in AI-powered validation so admins feel confident deploying AI-powered classifiers?
03/ THE PROCESS
Building trust in AI outputs to adopt feature and justify associated cost.
THE IDEA
AI-powered validation of the context of a file to weed out false positives at the root during classification.
WHAT
AI-powered validation layer
Ability to look at context and provide a verdict on sensitive or not sensitive content.
TRUST IS KEY
Trust in the AI-output
- Pain-point of false positives is reduced.
- Admins adopt and invest in AI-powered data classification
HOW
Explainability and Control
- Provide easy to understand reasoning for all validations.
- Ability to override, give feedback and improve AI results
- Provide metrics that help admins make the decision to use AI-powered data classification
01
Bringing transparency to how AI validates files.
HOW AI VALIDATES
LLM auto-generated natural language guidance from classifier configuration to perform validation actions.
Classifier configurations are complex
- LLMs are not able to understand the context directly from the classifier configurations
Auto-generates natural language guidance
- Understands context from name, description and the configuration to give itself a natural language ‘intent’ for the classifier.
Uses generated classifier guidance to validate
- The data classification system in Purview has a first pass on marking files for sensitivity
- AI validation runs on those files and removes sensitivity marker from the ones validated as false positives.
CONTEXT
Sensitive information types (SIT) are the most commonly used type of classifiers.
They are of 2 types:
- Out-of-box SITs
- Custom SITs
THE PROBLEM
For custom sensitive information types LLM could not produce reliable validation results.
How might we improve the classifier guidance for custom sensitive information types?
DECISION 1
- Classifier guidance included as part of Custom SIT configuration
- Classifier guidance to be auto-generated using available context
- Admins can review and edit classifier guidance to improve the validation results.
OUT-OF-BOX CLASSIFIERS
Created by Microsoft
- Established configuration
- Richer context for LLM
- Consistent validation results
CUSTOM CLASSIFIERS
Created by user
- Variable configuration
- Limited context, made for specific use case
- Less reliable validation results
CONSTRAINTS
- The review and edit of classifiers would be a user initiated process.
- No way for AI to rate the classifier guidance to signal to the user about the strength of the guidance.
- Time limitations. The experience needed to be locked in one week as it was a blocker for Alert Triage Agent. So the scope had to be scaled down considerably for the first release.
02
Building trust in the AI validation results
THE GOAL
How could we help the users make the decision to adopt AI-powered Sensitive information types?
THE FLOW
In a simulation environment
User chooses to test AI-powered SIT in a simulation environment
Data Classification System (DCS) classifies files
AI validates DCS classified files
Based on false positives removed AI-powered SIT accuracy score generated
User checks validation result
Publish
AI-powered SIT
DECISION 2
Quantify the validation results as an accuracy score for the classifier.
I chose to show this in two ways:
- Accuracy score for the classifiers.
- Comparison metrics of AI-powered SITs and the Base SIT
THE CHALLENGE
- Accuracy score would be the score for the validation on the sample data, not overall accuracy
- Accuracy score could possibly be low depending on the sample data and the classifier guidance.
USER TESTING INSIGHTS
“how does that translate into any business impact… like what’s the time to decision? Right: AI made a processing in 2 seconds per item, a human was doing 15 minutes per item – time savings is this much. Those values would definitely help.”
– HP
DECISION 3
Emphasis on what was de-classified rather than what was classified
Transparency into
- What was marked as false positive.
- Summary of reason why file was marked false positive.
- Detailed reasoning for further inspection
USER TESTING INSIGHTS
“I would want to see the actual match – the contextual match – in a column here… so that this is the file that it was found in, this is the AI decision, this is how confident it was. Here’s the match that it found, and here’s the reasoning why that match is not an actual positive.”
– The Hartford
03
Bringing AI-powered classification into Data Security Posture Management (DSPM)
ADAPTING TO BUSINESS PIVOT
- Leadership bundled AI classification into Data Security Posture Management, anticipating strong demand from top users and betting on AI.
- Shifted the product from reactive to proactive
- Concept of ‘Grading’ introduced
GRADING PROCESS
Proactive grading for high volume Out-of-box SITs
Data Classification System (DCS) classifies set number of files
AI-validation on DCS classified files
Grading Accuracy score generated
SIT Accuracy score higher than threshold
SIT Accuracy score lower than threshold
Automatic action:
Upgrade to AI-powered SIT
User action:
Re-run grading
User action:
Leave as Base SIT
User action:
Upgrade to AI-powered SIT
DECISION 4
Answers to ‘What does all the metrics mean? Where does it come from?’
Along with accuracy score, we added another metric for noise reduction.
We had received signals from the admins that though the metrics were helpful, they would want more details on how we arrived at those metrics
INSIGHT FROM USER TESTING
Metrics and accuracy scores draw high attention and seek explainability; “Accuracy %, & post publish % Accuracy difference” was met with scepticism. Users seek detailed definitions, calculation logics and transparency.
DECISION 5
Information needed to make the decision
CARD INFORMATION ARCHITECTURE
Following a lot of iterations and discussions we landed on a card showing only the most important information
- Accuracy score and matches
- Type of classifier
- Current phase of the classifier state
Japan My Number
Graded, Built-in
Accuracy score
90%
Total matches
22K
U.S. Passport Number
Upgrading…
U.S. Social Security Number - Employer Identification
AI-powered, Built-in
Accuracy score
92%
↑ 38%Total matches
30K
AI-POWERED SIT DETAILS PAGE
After taking into account the user feedback.
- More information on the time of the accuracy score
- Upfront explanation for the validations.
- More details on file level for matches
- Removed confidence level from previous version
MATCHED FILE DETAILS PAGE
After taking into account the user feedback.
- Contextual artifact added along with the validation and a reasoning from users
04/ ADDING IT ALL UP
05/ IMPACT
Solving for the most recurring pain-point for our customers
Feature in close testing with design partners like Nestlé among others.
One of the most anticipated capabilities for compliance admins.
06/ TAKEAWAYS
Work fast
This project taught me how to move quickly without losing the value of collaboration. With a tight timeline for user workshops and the added complexity of aligning with a different solution (DSPM), we had to make decisions, test ideas, and iterate quickly.
Working in a squad and across two geographies also reinforced how important clear communication and collaboration are when working in a follow-the-sun model.
I also learned to use AI tools more deliberately in the design process. Figma Make helped me quickly turn ideas into something tangible for discussion and exploration, while Copilot helped synthesise research data and accelerate the next round of iterations.
Fail fast
This project taught me the value of testing ideas early rather than waiting for them to be fully formed. With the growing demand for AI-powered experiences, we took concept screens to users at different levels of fidelity to understand expectations and validate our hypotheses early.
A “design with the user” approach helped us quickly identify what was working, iterate on the screens, and pivot when needed. It also meant that product strategy evolved alongside user feedback rather than in isolation.
Working this way by experimenting, learning, and adapting quickly, it felt much more like working in a start-up, and was a refreshing shift from more traditional product development.