Data Governance
Platform decisions with deep integration complexity, organizational change, and long-term data stakes.
This interactive experience is the shipped product itself — the same application code customers run in production, mounted read-only in your browser over a real sample journey. Not a video, not a mockup: because the demo and the product are one codebase, it can never drift from the real thing.
Inside this journey
-
Regulatory Outcome Discovery
Map stakeholder roles, required examiner timelines, in-scope data domains, and technical constraints the solution must satisfy.
Discovery Questions
Opening the conversation, how this reached your desk
- Tell me briefly how this regulatory issue first came to your team's attention
- When the examiner set the timeline, what deadline were you given to produce a full inventory and evidence of deletion or retention
- Who on your side is accountable for responding to the examiner, and who will be our day-to-day contact for discovery
- How many systems does your team estimate might hold customer personal information that we should prioritize in a first pass
Where the current inventory breaks under pressure
- What single missing artifact in your current inventory would make an examiner consider the response incomplete
- Walk me through the last time an examiner challenged your inventory, what information did they request beyond what you provided
- How often do such gaps occur across your systems
- Which data domains tend to be most problematic when tracing customer personal information in your estate
- If you could only guarantee evidence from one system within the examiner's timeline, which system would you pick and why
How this failure actually hits the business
- Who bears the regulatory and financial risk if the examiner rejects your response
- What financial or operational consequences followed the last negative exam finding
- Describe how much manual effort and how many FTE hours your team spends each month keeping the current inventory accurate
- Typically, what's the elapsed time from a regulator's request to a verifiable lineage report from your team
- If the examiner required proof of deletion across all downstream copies, what part of your environment would make you stop the project immediately
What's getting in the way right now
- Where do integration or access bottlenecks create the biggest risk of missing the examiner deadline
- Which teams own the connectors or APIs we would need to crawl, and how responsive are they to urgent access requests
- Do you have a central identity or access management process that can issue time-bound credentials for the seller to perform read-only discovery
- Tell me about any encryption or masking in place that might prevent automated column discovery without additional approvals
- Identify the single access gap, if unresolved, that would force you to extend the examiner's deadline
The alternatives you are weighing
- Name the main external solutions your team is actively evaluating and any internal option proposed to solve inventory and lineage gaps
- Under what conditions would you keep your current approach instead of switching to an external solution
- Has an internal team proposed to build automated discovery and lineage rather than buying a vendor, and if so what timeline and budget did they estimate
- Pick the single most important evaluation criterion when choosing between external vendors and an internal build
- Could any single missing capability from external vendors cause you to commit to an internal build instead
If this succeeds, what changes for you
- Imagine a 60-day proof-of-concept proved full discovery and lineage for your highest-risk domain, what would change in your next examiner interaction
- List the measurable acceptance criteria your team would require from a 60-day POC to sign a commercial agreement
- Identify the stakeholders who must approve POC success, and their individual success metrics or redlines
- Within your target timeline, what level of lineage depth do you expect the seller to produce for the pilot domain
- Assuming the POC meets the acceptance criteria, what internal obstacle could still prevent your procurement or legal team from signing the contract within 30 days
Realities that will block or speed implementation
- Name the non-negotiable technical or compliance constraints we must meet before beginning discovery
- Do you currently have service accounts or vetted connectors the seller can use for read-only crawling
- Point to the owner or team who controls API access for each critical system and whether escalation paths exist
- Is there an active legal or compliance review process that will need to clear a third-party discovery run before credentials are issued
- List any external vendor contractual terms or insurance requirements that would block discovery without higher approval
The technical path: connectors, latency, and non-invasiveness
- Given your current architecture, what integration approach would be unacceptable to your CTO
- Provide the top five source systems or storage endpoints we must crawl during the POC
- Are there encryption, tokenization, or masking layers that will prevent column-level inspection without additional approvals
- Estimate the minimum window of uninterrupted access you can provide during the 60-day POC for continuous crawling and lineage capture
- Point out the single integration or security concern that would stop your CTO from approving the POC
- Provide the expected SLA for discovery jobs in production and any maximum acceptable performance impact thresholds
What examiner-ready evidence must look like
- Describe the minimum evidence package the examiner will accept to close this finding
- Outline the stakeholders who must receive the evidence and their preferred format
- Within what timeframe after POC completion do you need the first examiner-ready package
- Would a demonstration of automated deletion certificates for downstream copies close the finding, or does the examiner require additional steps
If we proceed, what must happen next
- Assuming we proceed, what authorization and legal steps need to be finished before we can start the POC within your requested timeline
- Share the top three systems you want prioritized in the 60-day POC and briefly explain why each is critical
- Estimate the internal resource commitment you can allocate to the POC in FTEs and approximate weekly hours per FTE
- State the single event that would allow your procurement team to sign within 14 days after a successful POC
- When would you be ready to start a discovery run if all approvals are in place
-
POC Scope & Acceptance Criteria
Define the 60-day POC boundaries, target systems, measurable acceptance criteria for discovery accuracy, lineage depth, and policy enforcement.
Scope Configuration
- Automated Production Database Discovery
- Automated Data Warehouse Connector Deployment
- Cloud Object Storage Discovery and Cataloging
- Analytics and BI Platform Connector Deployment
- Streaming and Message Bus Connector Deployment
- Column-Level Lineage Reconstruction (SQL & Code)
- Non‑Invasive Pipeline Integration and Metadata Capture
- Live Data Catalog with Column Metadata
- Automated Retention and Deletion Enforcement
- Automated Data Masking and Pseudonymization
- Continuous Lineage and Schema Change Monitoring
- Regulatory Evidence Package Generation
Scope Questions
Automated Production Database Discovery
- Which production database engines store customer personal information (for example, core banking PostgreSQL, Oracle, or mainframe DB)?
- How many distinct database instances (by host:port) contain customer account numbers or personally identifiable information (PII) you want crawled during the 60-day POC?
- Do you have any read-only replicas or reporting clones that must be excluded to avoid double-counting records?
- List the schemas and example table names that hold KYC or customer contact fields we should prioritize (e.g., customers.customer_profile, loans.accounts).
- Provide the authentication method available for each database (username/password, certificate, cloud IAM role) and indicate where credentials will be stored.
Automated Data Warehouse Connector Deployment
- Identify the data warehouse platforms to connect for the POC (for example, Snowflake, Redshift, BigQuery) and the analytic schemas that host customer PII.
- Which warehouse schemas contain denormalized customer datasets used by reporting teams (for example, customer_ods, marts.customer_analytics)?
- Specify expected average daily query volume during discovery to estimate connector load on the warehouse.
- Confirm whether we can create a read-only service account with usage limits and list any approval steps required by your security team.
- Provide sample SQL or table access controls (grant statements) that our connector will need to discover column-level metadata.
Cloud Object Storage Discovery and Cataloging
- List the cloud object storage buckets that may contain customer documents or exported PII (for example, S3 prefixes or GCS buckets).
- Indicate file formats present (CSV, Parquet, JSON, Avro) and which prefixes hold extracts of account or transaction data.
- Where are retention policies that govern object storage exports of customer data documented (policy repo, compliance binder)?
- Do you use server-side encryption or customer-managed keys for buckets that contain PII?
- Attach an example path to a sensitive object and sample object metadata we can scan during the POC.
Analytics and BI Platform Connector Deployment
- Identify BI tools that host dashboards built from customer datasets (for example, Tableau workbooks, Looker explores) that need lineage tracing.
- How many dashboards, scheduled extracts, or data models reference customer PII that we should include in scope?
- Are dashboards embedding SQL-derived fields or referencing derived tables created by ETL jobs such as Airflow DAGs or Spark jobs?
- State the connectors that require OAuth or SSO and note whether admin consent is pre-approved.
- Specify whether BI extracts run on a schedule or are live queries and the typical extraction window that may impact POC performance.
Streaming and Message Bus Connector Deployment
- Name the streaming platforms that carry customer events (for example, Kafka topics, Kinesis streams) and specify the topics containing PII.
- Estimate the partition count and consumer groups for high-volume topics used for transaction-level data.
- Indicate whether streams undergo schema evolution (Avro or Protobuf) and whether a schema registry exists.
- State the retention window for each topic containing customer identifiers and whether compaction is enabled.
- Share connector network details including bootstrap servers, ports, and TLS requirements.
Column-Level Lineage Reconstruction (SQL & Code)
- Describe the primary ETL technologies that transform customer fields (for example, Airflow, dbt, Spark, stored procedures) and where transformation code is stored.
- Report the number of distinct transformation jobs that reference customer identifiers or PII and include representative job names.
- Enumerate SQL patterns commonly used to derive customer fields (joins, window functions, regex extraction) that impact lineage reconstruction.
- How will we access in-repo code artifacts (git URLs, access tokens) for code-based transformations to parse column mappings?
- What acceptance criteria will confirm the POC has reconstructed column-level lineage from production reporting tables back to original source columns for a sample set of 50 pipelines?
Non‑Invasive Pipeline Integration and Metadata Capture
- Explain any constraints that prohibit installing agents or modifying existing ETL pipelines in your core banking or fraud detection environments.
- Detail which metadata surfaces are available for passive capture (query logs, job run metadata, orchestration APIs) in your data platform.
- Are there compliance rules that prevent connecting to audit logs or query history from the fraud analytics cluster?
- Report any windows of change freeze during the POC (for example, month-end close) when pipeline metadata is unstable.
- Share owner contacts for pipeline teams and the escalation path if connector activity triggers operational alerts.
Live Data Catalog with Column Metadata
- Declare canonical identifiers that must be visible in the live catalog for examiners (for example, customer_id, ssn_hash, account_number).
- How frequently should column metadata refresh to capture schema drift in reporting tables used by compliance?
- Should lineage-annotated column descriptions include legal basis or retention policy references used by general counsel during examinations?
- Estimate expected column-level cardinality ranges for representative customer tables to size catalog storage.
- Who will own catalog curation and approve edits to column metadata during the POC?
Automated Retention and Deletion Enforcement
- Describe your current statutory retention schedules for customer records (for example, account statements: 7 years, transaction logs: 3 years) and where they are stored.
- Explain which systems currently perform deletion workflows and how deletions are documented for examiners (audit trail, deletion receipts).
- What automated thresholds would you accept for POC success (for example, automated deletions executed within 7 days of policy expiry for 95% of eligible objects)?
- Confirm whether deletions must propagate to downstream analytic copies, archived backups, and exported CSVs and name the archive storage locations.
- Who will approve retention rule changes and which legal artifact documents the retention schedule (policy ID, legal memo) that examiners expect?
Automated Data Masking and Pseudonymization
- Enumerate the PII categories that require masking or pseudonymization for analytics (for example, SSN, full name, birthdate, payment card data) and preferred masking techniques.
- Name downstream consumers that expect masked values versus true values for production reporting or modeling.
- Outline key rotation and key management requirements if pseudonymization relies on deterministic encryption and customer-managed keys.
- Clarify whether trial datasets used for model development require reversible pseudonymization under legal hold and which legal holds currently exist.
- Give an example of a table and column where masking must preserve analytic aggregation properties (for example, masked credit_score still supports average calculation).
Continuous Lineage and Schema Change Monitoring
- How quickly must the system detect schema changes for customer-facing reporting tables (for example, within 24 hours of deployment)?
- Select preferred notification channels for schema-change alerts to data owners (email, Slack, ticketing system) and indicate severity thresholds.
- Declare the expected retention period for historical lineage graph versions for examiner review (for example, 90 days, 7 years).
- Clarify whether scheduled schema migrations (for example, quarter-end normalization) should be excluded from alerting windows.
- Declare recipients for monthly lineage health reports and the KPIs they should contain (discovery accuracy, drift rate, missing lineage percent).
Regulatory Evidence Package Generation
- Outline regulator templates or evidence formats examiners expect (for example, CSV audit logs, PDF retention certificates, chain-of-custody manifests).
- What evidence artifact will validate automated discovery accuracy for the POC (for example, a side-by-side CSV of discovered tables versus ground-truth inventory) and who will sign off?
- Detail the file formats and signing requirements for producing examiner-ready export bundles and whether notarization is required.
- Clarify whether evidence packages must include hash chains or checksums for each exported dataset and which hashing algorithm is required.
- Supply the contact on your legal or compliance team who will accept the final evidence package and the delivery method (secure portal, SFTP, encrypted email).
-
Proof-of-Concept Evaluation
Run the agreed 60-day production-connected evaluation to verify automated discovery, column-level lineage, and enforcement against acceptance criteria.
- desired_state
- current_state
- stakeholders
- gaps
- success_criteria
- decision_readiness
- desired_state
- decision_readiness
- current_state
- success_criteria
- gaps
- stakeholders
- desired_state
- success_criteria
- stakeholders
- gaps
- current_state
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
- decision_readiness
-
Mutual Commit
Finalize commercial and legal terms, confirm data-access authorizations, and lock CTO integration requirements and SLAs before full rollout.
Agreement Modules
- Subscription Agreement
- Master Services Agreement (MSA)
- Statement of Work (SOW)
- Service Level Agreement (SLA)
- Data Processing Agreement (DPA)
- Financial Services Compliance Addendum (Data Residency & SOC2)
- Data Access Authorization & Credentialing Form
- Technical Integration Requirements & Acceptance Annex
- Subprocessor Disclosure & Approval
- Change Order Agreement
-
Deployment
Lock readiness facts and configuration values before execution begins.
-
Pre-Deployment Readiness
Confirm concrete readiness facts — environments, access, owners, and retention obligations — required before execution begins.
Pre-Deployment Questions
Environment and access
- Which environment types must the seller have read access to for this rollout? (select all that apply)
- Is production access already approved for the rollout start date? If not, provide the date read access will be granted or 'Already available' (this schedules our first execution window)
- Who is the technical owner authorized to provision access for the listed environments? Provide name and role (so the deployment team can coordinate provisioning)
Data and configuration
- Which data domain(s) are in-scope for initial crawling and validation? (select all that apply)
- Has the buyer decided the authoritative source-of-truth and field-mapping approach for the in-scope domain(s)?
- If 'Yes' above, who owns the source-of-truth and mapping (name and role)? If undecided, state who will finalize it (this owner signs off on schema alignment)
People and ownership
- Please list the primary named owners (name + role) for: deployment lead, security/compliance liaison, data owner for each in-scope domain, and legal/compliance approver (these people must be available for approvals during rollout)
- Do the named owners have authority to approve access changes and runbooks during the evaluation window?
Timing and constraints
- List any operational blackout windows, regulatory reporting dates, or other dates when work against production systems is prohibited (select all that apply)
- Are there binding retention, deletion, or export constraints that will affect testing (for example: no data exports, masking required, or deletion validations)? Select the option that best describes the requirement.
- Is there a confirmed go/no‑go decision date and an authorized decision‑maker for starting production crawling? If yes, state 'Confirmed' and name the decision‑maker; if no, state expected decision date or 'TBD'.
-
Configuration Details
Capture exact integration values the deployment needs — connectors, credentials, endpoints, performance constraints, and non-invasive integration requirements.
Configuration Details
Environments & Endpoints
- Enter the primary production environment name (enter the exact environment identifier the buyer uses; e.g., 'prod-us-east-1')
- Primary deployment region (Default: us-east-1 — change only if another cloud region is required)
Connectors & Sources (initial POC scope)
- Select the single primary source system category the deployment will connect to for the initial POC (choose one primary category)
- Enter the exact connector instance identifier or alias for the selected source (the integration name from the buyer's connection catalog). Do not paste secrets.
Authentication & Secrets Handoff
- Select the authentication method the buyer will provide for this connector (Default: Service account / integration user)
- Enter the non-secret credential identifier the buyer will supply for this connector (integration username, client ID, IAM role name, or certificate name). Do not paste secrets.
- Where will the buyer place the connector secrets for secure exchange at deployment kickoff? (select one)
Non‑invasive Integration & Feature Toggles
- Will the deployment require read-only, non-intrusive access to production systems? Default: Yes
- Select which feature modules to enable for this deployment (select one or more). Default expected for POC: Automated discovery, Column-level lineage, Policy enforcement
-
Deployment
Execute production crawling, lineage capture, and policy enforcement rollout with clear owners, sequencing, and monitoring checkpoints.
-
-
Success
Validate outcomes against acceptance criteria, produce examiner-ready evidence, and maintain a channel for issues and enhancement requests.
Success Reviews
- Go-live health check (weeks 1-4)
- First measurement review (weeks 4-10)
- Examiner evidence validation and handoff (approx. day 90)
- Operational enhancement and incident triage (monthly)
- Quarterly operational review (quarterly)
- Annual compliance readiness review (annual)
Issues & Enhancements
- Approve the prioritized backlog for the next quarter focused on closing remaining gaps.
- Reduce the count of open high-severity incidents and assign resolution windows for each.
- Agree the top 5 enhancement priorities for the next delivery cycle to support the acceptance criteria.
- Confirm the incident notification and evidence request process for regulators.
- Create fixes or workarounds for the top two high-severity incidents and schedule testing.
- Update the incident SLA tracker with measured mean time to resolve values and next milestones.
- Deliver a short incident root-cause report for each resolved enforcement failure.
- Quarterly metrics dashboard
- Confirm that catalog coverage and enforcement metrics are stable or improving toward the targets recorded in the POC Scope & Acceptance Criteria stage.
- Re-confirm deployment checklist and owners
- Validate retention and archive samples meet regulatory evidence requirements.
- Schedule remediation sprints for the top coverage gaps identified in the metrics dashboard.
- Publish the quarterly evidence sampling report and archive index.
- Update the operational runbook to reflect any process changes agreed in the meeting.
- Yearly metric trend analysis
- Confirm the organization can produce examiner-ready evidence within the mean time targets recorded in the POC Scope & Acceptance Criteria stage.
- Validate archive integrity and that the percentage of verified deletion requests meets compliance expectations.
- Agree the annual drill schedule and responsibilities for evidence retrieval testing.
- Run an annual evidence retrieval drill and publish the results with any remediation steps.
- Update retention and deletion configurations to reflect any regulatory or policy changes.
- Schedule the next annual compliance readiness review and interim sampling checkpoints.
- Confirm production crawling, lineage capture, and enforcement rollouts are running against the deployed connectors and environments.
- Document the top 5 open issues with remediation actions and target completion dates.
- Confirm schedule for the First measurement review meeting within the 4-10 week window.
- Publish a deployment health report capturing connector status, recent crawl results, and error logs.
- Log remediation tasks with resolution dates and evidence acceptance criteria references.
- Circulate user onboarding checklist and initial training completion summary.
- Present first-run metrics versus targets
- Agree a remediation plan with dates to bring discovery precision rate and column-level lineage coverage to the targets recorded in the POC Scope & Acceptance Criteria stage.
- Identify the top 3 technical causes for shortfalls and schedule fixes in the operational backlog.
- Confirm the exports and evidence artifacts required for the examiner-ready package and the delivery date for each.
- Implement prioritized connector or mapping fixes to improve discovery precision rate.
- Run a targeted re-crawl and lineage capture for the affected pipelines and publish results.
- Deliver the first evidence artifact pack for review prior to the Examiner Evidence Validation meeting.
- Restate acceptance criteria sources
- Deliver a complete examiner-ready evidence package that maps to the POC Scope & Acceptance Criteria artifacts and metrics.
- Confirm archive location and retention metadata for all evidence artifacts required for regulator review.
- Document the incumbent decommissioning status and next steps to prevent parallel usage.
- Publish the finalized examiner-ready evidence package to the agreed repository with retention metadata.
- Execute the agreed incumbent wind-down steps: archive or set read-only and record the retention decision.
- Document any remaining remediation items from evidence validation with target completion dates.
- Review open high-severity incidents
- Deployment and integration validation
- Present the examiner-ready evidence package
- Root cause analysis for metric gaps
- Persistent issues and systemic risks
- Archive integrity and sampling results
- Discuss mean time to resolve enforcement failures
- Verify metrics against evidence
- Regulatory change and retention obligations review
- Agree corrective actions and timelines
- Early adoption and usage signals
- Prioritize enhancement requests
- Backlog prioritization and resourcing
- Open issues and blockers
- Confirm evidence collection plan for final validation
- Plan for next-year evidence drills
- Archive and access plan for evidence
- Confirm communication plan for audit and incident notification
- Retention and evidence archive compliance check