How schemas are implemented
How MikroORM's schema entity option is implemented for Neo4j: where the implicit predicates are
injected, what they cost, and what deliberately stays unscoped. The user-facing guide is
Schemas and multi-tenancy.
mikro-orm-neo4j ≥ 0.3.0 · @mikro-orm/core 7.x · Neo4j ≥ 5.7
Summary
| Area | Before | After |
|---|---|---|
| Entity option | schema accepted by core, silently ignored by the driver — every entity in one flat graph | Fixed ({ schema: 'audit' }) and wildcard ({ schema: '*' }) both honoured |
| Node identity | MERGE keyed on the primary key | Keyed on (…primary key, __schema): the same id under two schemas is two nodes |
| Reads | Unscoped | find / findOne / count filter on the resolved schema implicitly |
| Writes | Unscoped | nativeUpdate / nativeDelete match within the schema; __schema is immutable |
| Relationships | Endpoints matched on their full primary key (C11) | …and on each endpoint's own resolved schema |
| Pivot entities | Endpoints matched on their full primary key | Endpoints additionally constrained by their nodes' schemas |
| Populate | Target matched by label | Target matched by the schema resolved from its own metadata |
| Hydration | — | The reserved __schema key is stripped; a user property named schema is not |
| Index generation | User declarations only | Composite identity constraint per entity, __schema-prefixed uniques/RANGE indexes, one global marker index |
| Migration | — | assignDefaultSchema() backfills nodes written before the entity opted in |
Isolation here is soft: everything lives in one Neo4j database and raw Cypher can still read across schemas. It is the graph analogue of SQL schemas, not of separate databases.
Storage model
A schema-aware node carries the resolved schema as a property and gains a marker label:
(:Invoice:SchemaNode { id: "INV-1", __schema: "tenant-1", total: 90 })
The property is the source of truth — merge keys, predicates and constraints all read it. The label exists only so cross-entity administration (drop a tenant, enumerate schemas) has one thing to match on, backed by a single global index.
The resolved schema is always written, including the default public. "Absent means default"
would force IS NULL predicates and make the composite index unusable.
Alternatives considered and rejected:
| Alternative | Why not |
|---|---|
Hub node (:SchemaNode { schema }) + [:IN_SCHEMA] edges | Supernode. Relationship creation locks both endpoints, so every write in a tenant serializes on the hub, and every query pays an extra hop. |
A dynamic label per schema (:Invoice:tenant_1) | Neo4j uniqueness constraints are single-label, so "unique per (Invoice, tenant)" is inexpressible. Labels also cannot be parameterized in Cypher. |
| Native multi-database | That is hard isolation — a different feature, reachable by pointing MikroORM at another database. |
One resolution choke point
Every schema decision in the driver goes through Neo4jDriver.resolveSchema, so
grep resolveSchema enumerates the complete surface where an implicit predicate can appear:
resolveSchema(meta, options) {
if (!Neo4jCypherBuilder.isSchemaAware(meta)) return undefined; // never opted in → do nothing
return this.getSchemaName(meta, options) ?? this.platform.getDefaultSchemaName();
}
getSchemaName is core's own chain — fixed entity schema → options.schema → options.parentSchema
→ configured default — so precedence matches the SQL drivers exactly. Neo4jPlatform names the
default public. Every consumer reads undefined as "emit nothing": no label, no property, no
predicate.
Before / after Cypher
Insert — the schema joins the merge key
-- before
MERGE (n:Invoice { id: $id })
SET n.total = $total
-- after
MERGE (n:Invoice:SchemaNode { id: $id, __schema: $schema })
SET n.total = $total
__schema is in the MERGE pattern and never in SET: that is what makes it immutable and what
makes two schemas holding one id two distinct nodes.
Read — an implicit predicate beside the user's filters
-- after
MATCH (this0:Invoice:SchemaNode)
WHERE (this0.__schema = $param0 AND this0.id = $param1)
RETURN this0 AS node
Update — scoped match, and __schema is never assigned
MATCH (this0:Invoice:SchemaNode)
WHERE (this0.__schema = $param0 AND this0.id = $param1)
SET this0.total = $param2
RETURN this0
Relationships — each endpoint answers for its own schema
-- before (post-C11): full primary key on both endpoints
MATCH (a) MATCH (b:Currency)
WHERE a.id = $sId AND b.id = $tId
MERGE (a)-[r:PRICED_IN]->(b)
-- after: identity now includes the schema, resolved per endpoint
MATCH (a) MATCH (b:Currency)
WHERE (a.id = $sId AND a.__schema = $sSchema)
AND (b.id = $tId AND b.__schema = $tSchema)
MERGE (a)-[r:PRICED_IN]->(b)
$sSchema and $tSchema are resolved independently, the target with the source's schema as
parentSchema. That is precisely what lets a wildcard tenant entity reference a fixed public
catalog entity: the tenant end resolves to tenant-1, the catalog end to public, and the catalog
node is not duplicated per tenant.
This is the schema form of the C11 leak. MATCH … MATCH … WHERE … MERGE is cartesian: without the
schema term, two tenants sharing a business id would each get the edge.
Populate — inline on the pattern, not in the outer WHERE
MATCH (this0:Invoice:SchemaNode)
WHERE this0.__schema = $param0
OPTIONAL MATCH (this0)-[this1:PRICED_IN]->(this2:Currency:SchemaNode { __schema: $param1 })
RETURN this0 AS node, this2 AS rel_currency
The target's scope is written into the pattern. In an outer WHERE it would filter the root
row away rather than leaving the relation unmatched, turning "no currency in this schema" into "no
invoice".
Generated indexes and constraints
Per schema-aware entity, ensureIndexes() emits:
CREATE CONSTRAINT `Invoice___schema_id_unique` IF NOT EXISTS
FOR (n:`Invoice`) REQUIRE (n.`__schema`, n.`id`) IS UNIQUE
The database is told the same thing the write path believes: identity is (schema, …pk). Without
it, two schemas holding one id would be a uniqueness violation rather than two nodes.
__schema comes first so the constraint's backing index also serves the plain
"everything in this schema" scans the driver emits implicitly. User-declared uniques and RANGE
indexes are prefixed the same way — SQL parity, since a unique is per-table-per-schema:
CREATE CONSTRAINT `Invoice___schema_email_unique` IF NOT EXISTS
FOR (n:`Invoice`) REQUIRE (n.`__schema`, n.`email`) IS UNIQUE
TEXT and POINT indexes are left alone: Neo4j accepts exactly one property on them, so prefixing would turn a correctly declared index into an error. FULLTEXT scores text rather than seeking, so a prefix would change what was declared.
One global index backs the marker label, emitted once however many entities opt in:
CREATE RANGE INDEX `SchemaNode___schema_idx` IF NOT EXISTS
FOR (n:`SchemaNode`) ON (n.`__schema`)
Relationship (pivot) entities are skipped: the property lives on nodes, and a pivot is scoped through its endpoints.
Neo4j ≥ 5.7 is a hard requirement
Composite uniqueness constraints landed in Neo4j 5.7 (Community included). Schema support needs
them, so ensureIndexes() does not degrade silently to a non-unique index — that would leave
identity unprotected while looking healthy. On an older server the failure is re-thrown with the
requirement spelled out. NODE KEY remains Enterprise-only and is not used.
Cost
The implicit predicate is an equality on the leading column of an index that exists for it.
tests/Neo4jSchemaSupport.test.ts runs EXPLAIN on the exact statement the driver emits for a
scoped find and asserts the plan contains an index seek and no NodeByLabelScan.
Migrating existing data
Enabling schema on a populated entity makes its pre-existing nodes invisible: they carry no
__schema, and every query now filters on one. Backfill before deploying the entity change:
const generator = orm.schema as Neo4jSchemaGenerator;
await generator.assignDefaultSchema(Invoice); // the entity's own resolved schema
await generator.assignDefaultSchema(Invoice, 'tenant-1'); // or an explicit one
which runs, in batches:
MATCH (n:`Invoice`) WHERE n.`__schema` IS NULL
CALL { WITH n SET n.`__schema` = $schema SET n:`SchemaNode` } IN TRANSACTIONS OF 10000 ROWS
CALL … IN TRANSACTIONS needs an implicit transaction, so this must not run inside
em.transactional. Rolling back is removing the schema option; the leftover property and label are
inert, and MATCH (n:SchemaNode) REMOVE n.__schema, n:SchemaNode clears them.
Limitations
- Soft isolation. One database, one graph. Raw Cypher reads across schemas by design. Physical separation means a different database or instance.
- Unscoped surfaces, matching how MikroORM treats raw SQL:
em.run(), virtual-entityexpressions, and the query builder'spattern()/call()composition APIs. The query builder'screate()/merge()are raw composition too and take no schema. - Opt-in per entity — a deliberate divergence from SQL. In SQL every table physically lives in a
schema, so
em.schemaaddresses any entity. Replicating that here would makeem.schemafilter entities whose existing nodes have no__schema, silently returning nothing. So the gate sits in front of every entry point:FindOptions.schema, fork schemas andwithSchema()are all ignored for an entity that never declared the option, which keeps its Cypher byte-identical. - Nodes cannot move between schemas through the update path;
__schemais part of identity. Use a deliberate migration. - Reserved names.
__schemaandSchemaNodeon schema-aware entities. - The denormalized scalar foreign key stored as a node property keeps only the first primary-key column and is ambiguous across schemas — pre-existing, see the C10/C11 appendix. The edge carries the truth.
Test coverage
tests/Neo4jSchemaSupport.test.ts runs against a real Neo4j
(Testcontainers) over four fixtures — one per resolution mode plus a control that never opted in:
- Writes — fixed schema stores its name and the marker label; a wildcard resolves from the
fork; the default
publicis written, never left absent. - Precedence — a fixed schema beats
em.schema;FindOptions.schemabeats the fork. - Identity — one id under two schemas is two nodes; re-persisting
(id, schema)updates in place. - Reads — find and count are scoped per fork;
orderBy/limit/offsetoperate inside the scope; the identity map keeps twins apart. - Updates & deletes — a write in
t1leaves thet2twin untouched, and noSETever names__schema. - Relationships — a
t1chunk links only to thet1document (edge count tot2is 0 — the schema form of C11); populate resolves the right endpoint. - Cross-schema — a wildcard tenant entity points at the single fixed
publiccatalog node. - Pivots — created and queried within a schema; endpoints never cross.
- Query builder — implicit scope in the built Cypher and in
execute();withSchema/ignoreSchemain any chaining order; a no-op on entities that never opted in. - Hydration —
__schemaabsent from entity data, a user property namedschemaintact. - Back-compat — an entity without the option produces byte-identical Cypher with and without a schema set, and explicit forcing on it changes nothing.
- Migration — legacy nodes invisible before the backfill, returned after it.
- Cost —
EXPLAINon the emitted scoped find seeks an index instead of scanning the label.
tests/unit/Neo4jSchemaStatements.test.ts covers the generated statements without a database:
the composite identity constraint (including composite primary keys), __schema-prefixed uniques
and RANGE indexes, TEXT left single-property, the marker index emitted exactly once, nothing at all
for entities that never opted in, and a full getCreateSchemaSQL() snapshot.