Data Access Control
Data access control decides what a person or agent can see, and enforces it at compile time - before any query reaches your warehouse. Personas, scopes, and policies resolve an allowed semantic subgraph during query planning; a request outside that subgraph fails to compile, so the data is never read. This page covers the policy types and how to define them.
Where BI-layer filters fall short
Traditional architectures mask data after the query runs. The warehouse reads the full row. The BI tool applies the masking view. The audit log records that the raw data was retrieved. This is the confused-deputy problem: the warehouse has no idea who the real user is. It only knows the service account that connected. Any agent or SQL client that bypasses the BI tool sees all rows. The masking layer is optional, not structural.
Colrows takes the opposite approach. Policy is part of the semantic graph. During query planning, the requester's persona resolves an allowed subgraph. Compilation occurs entirely within it. If a metric depends on a node outside that subgraph, resolution fails. There is no query to run.
Compile-time, not after-the-fact
In a traditional architecture, a query runs against the raw schema and a row-mask or column-redaction layer trims the result. The data is already touched, the audit trail begins after the fact, and the controls are only as good as the masking layer's coverage. Colrows takes a different approach: policy is part of the graph. During semantic binding, the requester's persona resolves an allowed subgraph; compilation occurs entirely within it. If a metric depends on a node outside that subgraph, resolution fails. There is no query to run.
How Colrows differs
| Dimension | BI-Layer Filters | Colrows Compile-Time |
|---|---|---|
| Execution point | Runtime (post-read) | Compile time (pre-query) |
| Auditability | Weak; log-dependent | Deterministic; full trace |
| SQL performance | High latency (filter after) | Optimized (predicate pushdown) |
| Agent ready | No (bypassed by direct query) | Yes (all query paths) |
The pieces
| Object | What it does |
|---|---|
| Persona | A first-class graph node representing a role. Holds policy bindings and scope. |
| Scope | The slice of the graph a request is allowed to traverse - global / datastore / persona / user. |
| Access policy | A named rule bound to a single datasource (optionally a schema), holding a set of same-typed permissions and the users and groups it applies to. |
| Permission | The smallest unit within a policy - which dataset and columns, under what row conditions. Two flavors today, fixed and regex, with attribute (ABAC) arriving in the next release. |
| Redaction policy | A separate control that masks a column's values (full / keep-first-N / keep-last-N) for specified users or groups. See Redaction Policies. |
Fixed permissions
The most granular flavor. A permission names an exact dataset (no regex), the columns to withhold via columnExclusion, and optional rowPredicates - SQL WHERE expressions validated when the policy is created. Only rows for which the predicates hold are returned, and the excluded columns never enter the plan. The schema is inherited from the enclosing policy.
"type": "FIXED",
"permissions": [
{
"dataset": "orders", // exact name; no regex
"columnExclusion": ["email", "phone"],
"rowPredicates": ["region = 'US'"]
}
]
Regex permissions
Column-level control across a range of datasets whose names match a regex - for catalogs whose object lists change frequently. Choose either allButTheseColumns (deny-list) or onlyTheseColumns (allow-list). Row predicates aren't supported here, because the dataset list isn't fixed.
"type": "REGEX",
"permissions": [
{
"dataset": "fact_.*", // every fact_* table in the schema
"allButTheseColumns": [".*_pii", ".*ssn.*"] // or onlyTheseColumns to allow-list
}
]
Column patterns are matched against the column names of the datasets the dataset regex selects.
Attribute permissions (ABAC) Coming soon
Attribute-based access control is on the roadmap and lands in the next release. Today you govern with fixed and regex permissions; the model below is a preview of the ABAC API so you can design for it now. Want early access? Talk to us.
Attribute-based access control will let a permission be a list of expressions combined with conjunctions (AND / OR, evaluated with normal precedence). Each expression tests an attribute of an entity - USER, USERGROUP, TABLE, or COLUMN - against a set of values. This is how you will write rules against tags and properties rather than table names that drift.
"type": "ATTRIBUTE",
"permissions": [
{
"expressions": [
{ "entity": "TABLE", "attribute": { "type": "PROPERTY", "propertyName": "region" },
"op": "EQ", "values": ["EMEA"] },
{ "entity": "COLUMN", "attribute": { "type": "PROPERTY", "propertyName": "classification" },
"op": "NOT_IN", "values": ["restricted"] }
],
"conjunctions": ["AND"]
}
]
Planned operators: IN, NOT_IN, EXCEPT, EQ, NOT_EQ, GT, LT, LIKE. An attribute will target either the entity's NAME or one of its PROPERTY values (by propertyName).
A complete policy
A policy wraps one or more same-typed permissions with the datasource it applies to and the users and groups it binds. A rolling time window or a region restriction isn't a separate policy type - express it as a rowPredicate (fixed) today, or as an attribute expression once ABAC ships.
{
"name": "us_sales_analysts",
"datasourceId": "<datasource-id>",
"schema": "analytics", // optional; only if the datasource has schemas/catalogs
"type": "FIXED",
"permissions": [
{
"dataset": "orders",
"columnExclusion": ["email", "phone"],
"rowPredicates": [
"region = 'US'",
"order_date >= current_date - interval '90' day" // rolling 90-day window
]
}
],
"users": ["jsmith"], // usernames
"groups": ["grp_sales_us"] // user-group ids
}
Every permission in one policy shares the same type. A permission's scope never transcends its datasource - a deliberate constraint so the reach of any rule stays deterministic and easy to reason about.
Zero-trust data access
The combination of persona scope, governed permissions, and compile-time enforcement gives you zero-trust by default:
- Default deny. A persona starts with no access; every grant is explicit.
- No implicit relationships. Joins must be proven on the graph; a persona without access to one side cannot use it as a join key.
- Policy as data. Policies are defined through the governed access API and stored in the semantic layer; every change is attributed and auditable.
- Audit trail by construction. Every executed query carries the persona, scope, policies evaluated, and the structural reasoning trace.
Masking trusts that every query runner remembered to apply the masking view, every BI tool refreshed the column list, and every ad-hoc analyst followed the policy. Compile-time governance assumes none of those things are true.
Manage via API
Access policies are managed in Administration → Access Policies, and on the platform API. Create, update, and delete require an admin role; listing and reads require any authenticated session.
| Operation | Endpoint |
|---|---|
| Create a policy (admin) | POST /api/access/create |
| Update a policy (admin) | PUT /api/access/update |
| List policies | GET /api/access/list |
| Get a policy | GET /api/access/list/{policyId} |
| Delete a policy (admin) | DELETE /api/access/delete/{policyId} |
POST /api/access/create takes the policy body above; PUT /api/access/update takes the same body plus a policyId. During semantic binding, Colrows resolves the policy against your graph - so authorized columns are accessible, excluded columns simply don't exist in the allowed subgraph, and out-of-scope rows are filtered - all before any SQL is generated. No masking layer needed, no secondary audit log required; authorization is part of the query plan itself.
Frequently Asked Questions
How does Colrows achieve row-level security without warehouse-native views?
Row-level security in Colrows is a structural property of the semantic graph. During query planning, policies are compiled into SQL predicates before the query is generated. The resulting SQL includes the row-level WHERE clauses built-in, not applied afterward. This means the warehouse never needs a separate view or masking layer.
Can Colrows policies be audited for compliance?
Yes. Every compiled query carries a complete audit trace: persona, scope, policies evaluated, nodes resolved, and the final SQL generated. This trace is deterministic and point-in-time reproducible, making it ideal for compliance audits (SOC2, HIPAA, GDPR). The audit log shows exactly what was authorized and why.
What happens when a policy is violated during query planning?
Compilation fails immediately. No query is generated. No data is read. The failure reason is returned to the caller, making it clear what permission was denied and why. This fail-safe design is what makes compile-time enforcement a zero-trust approach.
Best practices
- Tag columns now, so that when attribute (ABAC) permissions arrive you can write rules against those tags and properties instead of table names that drift.
- Keep policies focused - one datasource and one permission type per policy; layer finer conditions through row predicates today, and ABAC expressions once they ship.
- Prefer redaction (partial masking) over outright denial when downstream signals still need a stable join key.
- Review the audit trail weekly - the failures (compilation refused) are as important as the successes.
Learn more
For a deeper architectural perspective on why compile-time governance is necessary, see Data Authorization: Why Security Fails in the Semantic Layer. That post covers the confused-deputy problem in detail, contrasts it with traditional BI-layer approaches, and explains why deterministic execution is the only safe model for enterprise data.
These same policies bind to identity over every integration surface, including AI agents connected through the MCP integration - an agent's metadata:read or data:query scope resolves through the identical persona and predicate stack described above.
Access policies are enforced independently of who curates the underlying definitions. Editing a metric or business term in Catalog changes what a term means - it never grants access to data the requester's persona wasn't already authorized to see.