Collections and datasets
Explains where data lives and why the collection is the unit that holds it.
Every piece of data in D.Hub sits inside a collection. A collection is not a folder but a workspace shared by one purpose. The source tables, pipelines, ontology, and dashboards needed for project management live in the same collection, and permissions are organised at that level too.
Inside a collection, a table with actual rows and columns is a dataset. A dataset is not a single file but a managed object with a schema and a history, and other assets connect to it by reference.
Not a folder but a workspace
A folder only holds files. Where a file came from and who may open it are questions handled outside the folder.
A collection carries those questions with it. One collection is a set of assets that reference each other, the unit that permissions apply to, and the range a pipeline can pick its inputs and outputs from. For a pipeline to read a dataset and load it into an entity, both must live in the same collection. Splitting collections is less about tidying files than about drawing the boundary of what will be used together.
Datasets carry a contract
The schema you define when creating a dataset is not a display note but a contract to keep. Column names, data types, and nullability become the premise of every later connection.
That contract shows itself when data moves inside a collection. If join_dt in the participation records arrives as text while the target property expects a timestamp, something has to convert it. If the source may contain empty values while the target forbids nulls, the load fails. Treating a dataset as a contract rather than a file turns those failures from sudden errors into conditions you can check in advance.
What a collection boundary decides
Creating a collection also creates one ontology scope for it. Entities and relations do not appear as separate items in the collection tree, but each one always belongs to a collection. A collection boundary settles three things at once.
| What | How it is decided |
|---|---|
| Permissions | Who can view and edit the set is decided per collection |
| Meaning | One collection corresponds to one ontology scope |
| Combinable range | It defines which assets pipelines and dashboards can draw on |
Because all three share the same line, deciding how to split collections is both permission design and model design.
One collection or several
Keep data in one collection when the same people use it for the same purpose. If assets reference each other often but sit in different collections, every pipeline and ontology pays the cost of crossing that boundary.
Split when the audience differs. If a project aggregate the whole team reads and an employee source only HR may open share a collection, permissions have to be carved up again inside it. The criterion is not the kind or size of the data but who uses it together.