| name | umbp-add-backend |
| description | Add a new storage medium to UMBP's distributed mode as a MediumBackend (HBM, SSD, CXL, a remote object store, a fake for tests). Covers the two shapes a backend can take — a PageMemorySource for anything page-addressable, or a full MediumBackend for a medium whose bytes are not directly addressable — plus registration, the heartbeat/event contract, and the failure modes that only show up under the peer service. Use when the user wants to add or debug a storage tier/medium/backend in src/umbp/distributed, mentions MediumBackend, PageBackend, PageMemorySource or BackendRegistry, or asks why a new tier's puts or gets are not working. |
Adding a storage medium to UMBP distributed mode
A backend owns bytes and publishes descriptors for them. It does not move
them. Everything in this skill follows from that one rule, which is stated at
the top of src/umbp/include/umbp/distributed/peer/backend/medium_backend.h
and enforced by the type system since Phase 6 (Init receives a
MemoryRegistrar, which has no Submit on it).
If you find yourself wanting to copy bytes inside a backend, stop — that is the
transfer layer's job, and the sibling skill umbp-add-transfer-engine covers
it.
First decide which of the two shapes you need
This is the whole design decision, and getting it wrong costs a rewrite.
Is your medium page-addressable — can you hand out a raw pointer that the
transfer layer can read and write directly?
| Yes (DRAM, HBM, CXL, a pmem mapping) | No (SSD, S3, a network filesystem) |
|---|
| Implement | PageMemorySource — 5 methods | MediumBackend — 23 methods |
| Reuse | All of PageBackend | Nothing; but see "staging" below |
| Effort | ~100 lines | ~450 lines |
| Example | hbm_backend.{h,cpp} | ssd_backend.{h,cpp} |
Most media are the first case. PageBackend's slot lifecycle, bitmap
allocator, event outbox, reaper, read leases and copy pins never consult the
tier, so a second paged medium needs no second copy of any of it.
Shape 1: a paged medium (implement PageMemorySource)
Five methods, in page_backend.h:
class PageMemorySource {
virtual bool Allocate(const std::vector<uint64_t>& sizes, std::vector<Buffer>* out) = 0;
virtual void Release() = 0;
virtual mori::io::MemoryLocationType LocationType() const = 0;
virtual int Device() const = 0;
virtual const char* Name() const = 0;
};
LocationType() and Device() are the facts a descriptor cannot recover.
They are the reason this interface exists: PageBackend mirrors them into every
TransferRef it publishes, and that is what makes a local transfer against your
medium select the right engine. Get them wrong and nothing fails at build or
init time — a transfer just silently picks the wrong engine, or no engine at
all.
The recipe
-
Write the source. Put it in its own file if it needs a new dependency
(hbm_backend.cpp exists so HIP stays out of page_backend.cpp).
bool MyPageMemorySource::Allocate(const std::vector<uint64_t>& sizes,
std::vector<Buffer>* out) {
std::vector<MyHandle> taken;
std::vector<Buffer> staged;
for (uint64_t size : sizes) {
if (size == 0) continue;
MyHandle h = MyAlloc(size);
if (!h.valid()) {
for (auto& t : taken) MyFree(t);
return false;
}
staged.push_back(Buffer{h.ptr, h.usable_size});
taken.push_back(h);
}
handles_.insert(handles_.end(), taken.begin(), taken.end());
out->insert(out->end(), staged.begin(), staged.end());
return true;
}
Report the usable size in Buffer::size, not the requested one — host
hugepage rounding makes the extra genuinely allocatable, and PageBackend
publishes what you report.
-
Add a factory returning std::unique_ptr<MediumBackend>, next to
MakePageBackend / MakeHbmBackend. It must return the interface: Phase 5
Rule A says only PoolClient::Init may name a concrete backend.
std::unique_ptr<MediumBackend> MakeMyBackend(uint64_t page_size, ...) {
return std::make_unique<PageBackend>(TierType::MY_TIER, page_size,
std::make_unique<MyPageMemorySource>(...),
std::move(buffer_sizes), pending_ttl,
read_lease_ttl);
}
-
Add a config struct in distributed/config.h holding your medium's own
knobs — no enabled flag: PoolClientConfig::medium is the single selector.
Do not extend DramOwnershipConfig — hugepages/NUMA/prefault are
meaningless for HBM, and a device ordinal is meaningless for host memory.
Each medium brings its own knobs; that asymmetry is why the seam is a class
and not an options struct.
-
Add a case to the medium switch in PoolClient::Init. A node registers
exactly one backend, chosen by config_.medium, so you add a case rather
than another if. This is the only file that changes outside your own:
case TierType::MY_TIER: {
backend = MakeMyBackend(page_size, config_.my, ...);
break;
}
The shared Init / Register / error handling below the switch is written
once for every medium.
Why one medium, not a tier stack. The routing plane does not tier: master
treats every advertised tier as an equally valid put target (Phase 4 deleted
the hardcoded tier orders), so a node registering two backends mirrors
across them rather than promoting/demoting between them. Heterogeneity comes
from different nodes picking different media. If you find yourself
wanting two live backends on one node, you are asking for a local tiering
policy that does not exist yet — that is a routing-plane change, not a
backend one.
-
Add the tier to TierType (types.h) if it is genuinely new. HBM=1, DRAM=2, SSD=3 already exist. BackendRegistry is a map<TierType, ...>, so
one instance == one medium and the enum value is the identity. Also add a
matching UMBPMedium value in common/config.h and map it in ToTierType,
or no user-facing config can select your medium.
That is the whole change. Routing, the peer service, the heartbeat and the batch
executors were all written against BackendRegistry and need no edit.
Metrics: do not write any
PoolClient::Init wraps whatever the medium switch produced in
InstrumentedBackend before registering it, and that decorator derives the
whole generic series — operations by outcome, bytes committed / resolved /
freed, batch depth, time spent inside the medium — from the MediumBackend
calls themselves. Your backend is charted the moment it is composed in, under
tier=<your tier>, in the panels of
examples/monitoring/grafana/dashboards/umbp_backends.json that already exist.
There is no metric to register, no dashboard to edit, and no counter to add to
your backend.
Override SampleMetrics() (from MetricSource) only for state the interface
cannot show from outside — what the device itself did, how full an internal
arena is. When you do, obey the one rule that keeps the single dashboard
working: publish under the generic name and put your specifics in a label.
std::vector<MetricSample> MyBackend::SampleMetrics() const {
return {MetricSample{MORI_UMBP_METRIC_BACKEND_MEDIUM_EVENTS_TOTAL,
MORI_UMBP_METRIC_BACKEND_MEDIUM_EVENTS_TOTAL_HELP,
{{"event", "device_read_error"}},
device_read_errors_.load(std::memory_order_relaxed)},
MetricSample{MORI_UMBP_METRIC_BACKEND_MEDIUM_STATE,
MORI_UMBP_METRIC_BACKEND_MEDIUM_STATE_HELP,
{{"state", "queue_depth"}},
depth, MetricKind::kGauge}};
}
A metric named after your medium (mori_umbp_myssd_reads_total) compiles and
scrapes fine, and is still wrong: it needs its own panel, which is how UMBP
ended up with a dashboard per medium and an SSD dashboard wired to counters
that lost their publisher. tier= and backend= are stamped by the publisher —
do not set them. Counter values must be monotonic; gauges may move either way.
See umbp/distributed/metrics/component_metrics.h.
Shape 2: a medium whose bytes are not addressable
If you cannot hand out a pointer, you have two options, and the cheap one is
almost always right.
Stage through registered host memory (what SsdBackend does)
Publish an ordinary registered host DRAM arena as your buffer, and move bytes
between it and your device inside the backend:
BatchAllocate reserves a staging page and publishes it. The writer RDMAs
into it knowing nothing about your medium.
BatchCommit spills that page to your device and returns the page to the
arena. The page is borrowed for the write, not the key's home.
BatchResolve fills a staging page from your device and publishes it under a
read lease; a reaper reclaims the page when the lease expires.
The cost is one host copy per side. The benefit is that your medium reaches the
data plane with zero new transfer-layer concepts — no new TransferRef
kind, no new engine, no chaining — and a remote peer reading from your node sees
a perfectly ordinary registered buffer and needs no code at all.
Add a real endpoint kind (bigger; not yet done in tree)
transfer_engine.h reserves this: a FileRef/ObjectRef plus a kind tag, a
new engine, and chaining in CompositeTransferEngine for the remote reader
(device → bounce → wire), which that class explicitly does not implement. Take
this path only when you need zero-copy/GDS; staging does not block it.
Things that bite in shape 2
- Exhaustion has no honest encoding. A
Resolve that cannot get a staging
page must report found=false, which makes the client exclude your node and
retry elsewhere — wrong, because your node does hold the key.
medium_backend.h records that the "not ready, retry here" state was proposed
and rejected. Size the arena for read concurrency, and pin the behavior in a
test so a future control-plane fix changes it deliberately.
- Contiguity.
PeerSsdManager::PrepareRead takes one (ptr, capacity), not
a scatter list, so a staged key must fit one contiguous page. SsdBackend
therefore allocates one buffer of staging_pages * page_size and refuses
keys larger than a page. In distributed mode master's page_size is the KV
block size, so 1 key == 1 page is the normal case.
- Never hold the arena lock across IO. Take the slot out under the lock,
release it, do the spill/fill, then re-lock to return the page. Holding it
stalls every concurrent
Allocate and Resolve.
The contracts that are easy to get wrong
These are the ones with no compile-time protection.
Events go in ONE bundle under ONE seq. The heartbeat concatenates every
backend's events (DrainAllBackends). Never emit one bundle or one seq per
medium — that breaks the ack / seq-gap full-sync recovery.
SnapshotOwnedKeysForFullSync must clear the outbox in the SAME critical
section as the snapshot. Two separate locks drop events committed in between.
SnapshotOwnedKeys is const and must not mutate; the full-sync variant is
not.
kFailedNoSpace vs kFailed. kFailedNoSpace means "medium exhausted,
retry on another peer". Use it only when another peer could plausibly succeed. A
key too large for your page size is kFailed — no peer would do better, and the
writer must not keep hunting.
Evict returns one result per key, in request order. The peer service sums
freed bytes for a key mirrored across media and relies on the positional match.
Return bytes_freed = 0 (not an error) for unknown, already-freed, or protected
keys; master retries protected ones next round.
BatchAbort is idempotent — a slot already reaped or never seen still
reports true.
Shutdown must tolerate a failed Init and must not run concurrently with
anything else. Deregister before releasing memory: the registrar may still
hold an MR over those pages.
Init is idempotent, and a backend is fully live when it returns — start
your own reaper thread there, so no caller has to know you have one.
Allocation is gated between ClearLocal() and ClearFullSyncAcked(). No
new owned key may appear in that window.
Threading: every method may be called from the peer service's gRPC handler
threads and the heartbeat thread concurrently.
SetAutoFlushHook's callback runs under your lock — it must be cheap and
must not re-enter the backend. Signal the heartbeat thread and return.
Testing
Two levels, both worth having:
-
Without the transfer layer — drive BatchAllocate / BatchCommit /
BatchResolve / Evict directly and assert the bookkeeping. A
LocalOnlyRegistrar (returns TransferRef::HostBytes, counts
register/deregister calls) is all the MemoryRegistrar you need; it is also
exactly what CompositeTransferEngine degrades to on a node with no RDMA.
-
Through a real CompositeTransferEngine — a Put and a Get moving actual
bytes. This is what catches a wrong LocationType()/Device(), because a
mislabeled endpoint selects the wrong engine and the copy either fails or
silently does nothing.
Assert the registration contract explicitly:
EXPECT_EQ(registrar.last_loc, mori::io::MemoryLocationType::GPU);
EXPECT_EQ(backend->BufferRef(0).device, 0);
For a staging backend, the highest-value test is that Commit returns the
page: configure 2 staging pages, do 10 sequential puts, assert all succeed. A
leaked page wedges the backend after staging_pages puts and nothing else
catches it.
Skip rather than fail when hardware is absent (if (!HaveGpu()) GTEST_SKIP()),
so the suite stays runnable on a CPU-only box.
Register the test in tests/cpp/umbp/distributed/CMakeLists.txt with
add_test(NAME ... COMMAND ...), not gtest_discover_tests — the file
explains why (discovery bakes in build-time cmake paths and goes red when the
image's ctest runs it).
Building and running
The umbp C++ build needs protoc and grpc_cpp_plugin, which live in the mori
Docker image, not on the bare host:
docker exec <mori-container> bash -c "cd <repo> && \
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DBUILD_UMBP=ON -DBUILD_TESTS=ON -DGPU_TARGETS=gfx950 \
-DCMAKE_CXX_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \
-DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ && \
ninja -C build && cd build && ctest -E '^cco_' -LE integration"
Exclude ^cco_ — those are unrelated GPU collective tests and some hang without
a full fabric. -LE integration skips the tests needing real RDMA.
Reference
medium_backend.h — the interface, and a list of things deliberately not
on it (with reasons, so they are not re-proposed)
page_backend.h — PageMemorySource, HostPageMemorySource, PageBackend
hbm_backend.{h,cpp} — the minimal paged medium (shape 1)
ssd_backend.{h,cpp} — the staged, non-addressable medium (shape 2)
doc/design-backend-agnostic-refactor.md — §2 the descriptor/pointer rule,
§3 the three-component split, §8 what none of it fixes