Skip to main content

How schemas are implemented

How MikroORM's schema entity option is implemented for Neo4j: where the implicit predicates are injected, what they cost, and what deliberately stays unscoped. The user-facing guide is Schemas and multi-tenancy.

Applies to

mikro-orm-neo4j ≥ 0.3.0 · @mikro-orm/core 7.x · Neo4j ≥ 5.7

Summary​

AreaBeforeAfter
Entity optionschema accepted by core, silently ignored by the driver — every entity in one flat graphFixed ({ schema: 'audit' }) and wildcard ({ schema: '*' }) both honoured
Node identityMERGE keyed on the primary keyKeyed on (…primary key, __schema): the same id under two schemas is two nodes
ReadsUnscopedfind / findOne / count filter on the resolved schema implicitly
WritesUnscopednativeUpdate / nativeDelete match within the schema; __schema is immutable
RelationshipsEndpoints matched on their full primary key (C11)…and on each endpoint's own resolved schema
Pivot entitiesEndpoints matched on their full primary keyEndpoints additionally constrained by their nodes' schemas
PopulateTarget matched by labelTarget matched by the schema resolved from its own metadata
Hydration—The reserved __schema key is stripped; a user property named schema is not
Index generationUser declarations onlyComposite identity constraint per entity, __schema-prefixed uniques/RANGE indexes, one global marker index
Migration—assignDefaultSchema() backfills nodes written before the entity opted in

Isolation here is soft: everything lives in one Neo4j database and raw Cypher can still read across schemas. It is the graph analogue of SQL schemas, not of separate databases.


Storage model​

A schema-aware node carries the resolved schema as a property and gains a marker label:

(:Invoice:SchemaNode { id: "INV-1", __schema: "tenant-1", total: 90 })

The property is the source of truth — merge keys, predicates and constraints all read it. The label exists only so cross-entity administration (drop a tenant, enumerate schemas) has one thing to match on, backed by a single global index.

The resolved schema is always written, including the default public. "Absent means default" would force IS NULL predicates and make the composite index unusable.

Alternatives considered and rejected:

AlternativeWhy not
Hub node (:SchemaNode { schema }) + [:IN_SCHEMA] edgesSupernode. Relationship creation locks both endpoints, so every write in a tenant serializes on the hub, and every query pays an extra hop.
A dynamic label per schema (:Invoice:tenant_1)Neo4j uniqueness constraints are single-label, so "unique per (Invoice, tenant)" is inexpressible. Labels also cannot be parameterized in Cypher.
Native multi-databaseThat is hard isolation — a different feature, reachable by pointing MikroORM at another database.

One resolution choke point​

Every schema decision in the driver goes through Neo4jDriver.resolveSchema, so grep resolveSchema enumerates the complete surface where an implicit predicate can appear:

resolveSchema(meta, options) {
if (!Neo4jCypherBuilder.isSchemaAware(meta)) return undefined; // never opted in → do nothing
return this.getSchemaName(meta, options) ?? this.platform.getDefaultSchemaName();
}

getSchemaName is core's own chain — fixed entity schema → options.schema → options.parentSchema → configured default — so precedence matches the SQL drivers exactly. Neo4jPlatform names the default public. Every consumer reads undefined as "emit nothing": no label, no property, no predicate.


Before / after Cypher​

Insert — the schema joins the merge key​

-- before
MERGE (n:Invoice { id: $id })
SET n.total = $total

-- after
MERGE (n:Invoice:SchemaNode { id: $id, __schema: $schema })
SET n.total = $total

__schema is in the MERGE pattern and never in SET: that is what makes it immutable and what makes two schemas holding one id two distinct nodes.

Read — an implicit predicate beside the user's filters​

-- after
MATCH (this0:Invoice:SchemaNode)
WHERE (this0.__schema = $param0 AND this0.id = $param1)
RETURN this0 AS node

Update — scoped match, and __schema is never assigned​

MATCH (this0:Invoice:SchemaNode)
WHERE (this0.__schema = $param0 AND this0.id = $param1)
SET this0.total = $param2
RETURN this0

Relationships — each endpoint answers for its own schema​

-- before (post-C11): full primary key on both endpoints
MATCH (a) MATCH (b:Currency)
WHERE a.id = $sId AND b.id = $tId
MERGE (a)-[r:PRICED_IN]->(b)

-- after: identity now includes the schema, resolved per endpoint
MATCH (a) MATCH (b:Currency)
WHERE (a.id = $sId AND a.__schema = $sSchema)
AND (b.id = $tId AND b.__schema = $tSchema)
MERGE (a)-[r:PRICED_IN]->(b)

$sSchema and $tSchema are resolved independently, the target with the source's schema as parentSchema. That is precisely what lets a wildcard tenant entity reference a fixed public catalog entity: the tenant end resolves to tenant-1, the catalog end to public, and the catalog node is not duplicated per tenant.

This is the schema form of the C11 leak. MATCH … MATCH … WHERE … MERGE is cartesian: without the schema term, two tenants sharing a business id would each get the edge.

Populate — inline on the pattern, not in the outer WHERE​

MATCH (this0:Invoice:SchemaNode)
WHERE this0.__schema = $param0
OPTIONAL MATCH (this0)-[this1:PRICED_IN]->(this2:Currency:SchemaNode { __schema: $param1 })
RETURN this0 AS node, this2 AS rel_currency

The target's scope is written into the pattern. In an outer WHERE it would filter the root row away rather than leaving the relation unmatched, turning "no currency in this schema" into "no invoice".


Generated indexes and constraints​

Per schema-aware entity, ensureIndexes() emits:

CREATE CONSTRAINT `Invoice___schema_id_unique` IF NOT EXISTS
FOR (n:`Invoice`) REQUIRE (n.`__schema`, n.`id`) IS UNIQUE

The database is told the same thing the write path believes: identity is (schema, …pk). Without it, two schemas holding one id would be a uniqueness violation rather than two nodes.

__schema comes first so the constraint's backing index also serves the plain "everything in this schema" scans the driver emits implicitly. User-declared uniques and RANGE indexes are prefixed the same way — SQL parity, since a unique is per-table-per-schema:

CREATE CONSTRAINT `Invoice___schema_email_unique` IF NOT EXISTS
FOR (n:`Invoice`) REQUIRE (n.`__schema`, n.`email`) IS UNIQUE

TEXT and POINT indexes are left alone: Neo4j accepts exactly one property on them, so prefixing would turn a correctly declared index into an error. FULLTEXT scores text rather than seeking, so a prefix would change what was declared.

One global index backs the marker label, emitted once however many entities opt in:

CREATE RANGE INDEX `SchemaNode___schema_idx` IF NOT EXISTS
FOR (n:`SchemaNode`) ON (n.`__schema`)

Relationship (pivot) entities are skipped: the property lives on nodes, and a pivot is scoped through its endpoints.

Neo4j ≥ 5.7 is a hard requirement​

Composite uniqueness constraints landed in Neo4j 5.7 (Community included). Schema support needs them, so ensureIndexes() does not degrade silently to a non-unique index — that would leave identity unprotected while looking healthy. On an older server the failure is re-thrown with the requirement spelled out. NODE KEY remains Enterprise-only and is not used.

Cost​

The implicit predicate is an equality on the leading column of an index that exists for it. tests/Neo4jSchemaSupport.test.ts runs EXPLAIN on the exact statement the driver emits for a scoped find and asserts the plan contains an index seek and no NodeByLabelScan.


Migrating existing data​

Enabling schema on a populated entity makes its pre-existing nodes invisible: they carry no __schema, and every query now filters on one. Backfill before deploying the entity change:

const generator = orm.schema as Neo4jSchemaGenerator;
await generator.assignDefaultSchema(Invoice); // the entity's own resolved schema
await generator.assignDefaultSchema(Invoice, 'tenant-1'); // or an explicit one

which runs, in batches:

MATCH (n:`Invoice`) WHERE n.`__schema` IS NULL
CALL { WITH n SET n.`__schema` = $schema SET n:`SchemaNode` } IN TRANSACTIONS OF 10000 ROWS

CALL … IN TRANSACTIONS needs an implicit transaction, so this must not run inside em.transactional. Rolling back is removing the schema option; the leftover property and label are inert, and MATCH (n:SchemaNode) REMOVE n.__schema, n:SchemaNode clears them.


Limitations​

  • Soft isolation. One database, one graph. Raw Cypher reads across schemas by design. Physical separation means a different database or instance.
  • Unscoped surfaces, matching how MikroORM treats raw SQL: em.run(), virtual-entity expressions, and the query builder's pattern() / call() composition APIs. The query builder's create() / merge() are raw composition too and take no schema.
  • Opt-in per entity — a deliberate divergence from SQL. In SQL every table physically lives in a schema, so em.schema addresses any entity. Replicating that here would make em.schema filter entities whose existing nodes have no __schema, silently returning nothing. So the gate sits in front of every entry point: FindOptions.schema, fork schemas and withSchema() are all ignored for an entity that never declared the option, which keeps its Cypher byte-identical.
  • Nodes cannot move between schemas through the update path; __schema is part of identity. Use a deliberate migration.
  • Reserved names. __schema and SchemaNode on schema-aware entities.
  • The denormalized scalar foreign key stored as a node property keeps only the first primary-key column and is ambiguous across schemas — pre-existing, see the C10/C11 appendix. The edge carries the truth.

Test coverage​

tests/Neo4jSchemaSupport.test.ts runs against a real Neo4j (Testcontainers) over four fixtures — one per resolution mode plus a control that never opted in:

  1. Writes — fixed schema stores its name and the marker label; a wildcard resolves from the fork; the default public is written, never left absent.
  2. Precedence — a fixed schema beats em.schema; FindOptions.schema beats the fork.
  3. Identity — one id under two schemas is two nodes; re-persisting (id, schema) updates in place.
  4. Reads — find and count are scoped per fork; orderBy / limit / offset operate inside the scope; the identity map keeps twins apart.
  5. Updates & deletes — a write in t1 leaves the t2 twin untouched, and no SET ever names __schema.
  6. Relationships — a t1 chunk links only to the t1 document (edge count to t2 is 0 — the schema form of C11); populate resolves the right endpoint.
  7. Cross-schema — a wildcard tenant entity points at the single fixed public catalog node.
  8. Pivots — created and queried within a schema; endpoints never cross.
  9. Query builder — implicit scope in the built Cypher and in execute(); withSchema / ignoreSchema in any chaining order; a no-op on entities that never opted in.
  10. Hydration — __schema absent from entity data, a user property named schema intact.
  11. Back-compat — an entity without the option produces byte-identical Cypher with and without a schema set, and explicit forcing on it changes nothing.
  12. Migration — legacy nodes invisible before the backfill, returned after it.
  13. Cost — EXPLAIN on the emitted scoped find seeks an index instead of scanning the label.

tests/unit/Neo4jSchemaStatements.test.ts covers the generated statements without a database: the composite identity constraint (including composite primary keys), __schema-prefixed uniques and RANGE indexes, TEXT left single-property, the marker index emitted exactly once, nothing at all for entities that never opted in, and a full getCreateSchemaSQL() snapshot.