| title | Multiple Detection Sources |
|---|---|
| group | Recipes |
| summary | Compose model predictions, draft annotations, or other app-owned detection streams over one media item. |
Use detections.sources when one media item needs more than one semantic
detection stream.
Common examples:
- model predictions plus app-owned draft annotations;
- accepted annotations plus the shape a user is currently drawing;
- output from two model versions for side-by-side review;
- persisted detections plus short-lived reviewer notes or correction overlays.
The library does not know those business meanings. It only composes sources, preserves source provenance, orders detections deterministically, and lets each source override box, mask, polygon, polyline, keypoint, and label presentation.
Do not use multiple sources when you only need per-class colors, confidence
filtering, or a different label format. Those are normal style concerns and are
better handled with BaseBoxStyle, BaseMaskStyle, BaseLabelStyle, or custom
style resolvers.
One media session can read many detection sources, but the renderer still sees
one active semantic DetectionFrame.
predictions source ┐
draft source ├─ composed hot frame ─ prepared artifacts ─ Pixi layers
review source ┘
Copied detections receive:
sourceId: the source entry that produced the detection;sourceDetectionIndex: the detection index inside that source frame before composition.
Those fields are provenance, not workflow state. Your app can decide that
"draft" means “unsaved human edits,” but supervision-js treats it as just
another ordered source.
This example renders model predictions as the global default source and renders draft annotations above them with a different box/label style and no masks.
import {
BaseBoxStyle,
BaseLabelStyle,
BaseMaskStyle,
BoxShape,
annotationRenderers,
createMediaSession,
} from "supervision";
const PREDICTIONS_SOURCE_ID = "predictions";
const DRAFT_SOURCE_ID = "draft";
const session = await createMediaSession({
container,
media,
detections: {
sources: [
{
appendable: {
datasetId: "video-123-predictions",
},
id: PREDICTIONS_SOURCE_ID,
requiredForPlayback: true,
},
{
appendable: {
datasetId: "video-123-drafts",
},
id: DRAFT_SOURCE_ID,
order: 10,
presentation: {
boxStyle: new BaseBoxStyle({
cornerRadius: 8,
fill: { alpha: 0.16, color: 0xf59e0b },
shape: BoxShape.RoundedRect,
stroke: { alpha: 1, color: 0xfbbf24, width: 4 },
}),
labelStyle: new BaseLabelStyle({
background: { alpha: 0.85, color: 0x78350f },
text: (detection) => {
const label = detection.className ?? "object";
const confidence =
detection.confidence === undefined
? ""
: ` ${Math.round(detection.confidence * 100)}%`;
return `Draft ${label}${confidence}`;
},
}),
maskStyle: null,
},
requiredForPlayback: false,
},
],
},
presentation: {
renderers: [
annotationRenderers.box({ style: new BaseBoxStyle() }),
annotationRenderers.label({
style: new BaseLabelStyle({ includeConfidence: true }),
}),
annotationRenderers.mask({
style: new BaseMaskStyle({ opacity: 0.65 }),
}),
],
},
});
await session.appendDetectionFrames(predictionFrames, {
sourceId: PREDICTIONS_SOURCE_ID,
});
await session.replaceDetectionFrames(draftFrames, {
sourceId: DRAFT_SOURCE_ID,
});The prediction source is required for playback, so a playback gate can wait for prediction coverage. The draft source is optional, so a missing or empty draft window does not block playback.
order controls draw order. Lower sources compose first. Higher sources render
later and appear on top.
When a session owns more than one appendable source, writes need a sourceId.
await session.appendDetectionFrames(predictionFrames, {
sourceId: PREDICTIONS_SOURCE_ID,
});
await session.replaceDetectionFrames([currentDraftFrame], {
sourceId: DRAFT_SOURCE_ID,
});
await session.clearDetectionFrames({
sourceId: DRAFT_SOURCE_ID,
});Appending, replacing, or clearing one source does not mutate the others. The session composes them again when the hot window refreshes.
The top-level presentation.renderers list selects the global annotation
renderers and supplies their styles.
A source-level presentation can override:
boxStylemaskStylepolygonStylepolylineStylekeypointStylelabelStyle
For each selected renderer:
undefinedfalls back to the global style;nulldisables that layer for detections from that source;- a style object overrides the global style for that source.
Interaction and focus presentation remain global. If they need source-aware
behavior, branch on detection.sourceId inside the style resolver.
The composed frame still reaches the renderer as one semantic DetectionFrame,
so buffering, render preparation, picking, focus, and playback synchronization
continue through the same engine path.
Most apps should use createMediaSession({ detections: { sources } }). Use
createCompositeDetectionFrameSource() directly only when you already manage a
lower-level renderer or need to test composition outside a media session.
import { createCompositeDetectionFrameSource } from "supervision";
const source = createCompositeDetectionFrameSource({
sources: [
{ frames: predictionFrames, id: PREDICTIONS_SOURCE_ID },
{ frames: draftFrames, id: DRAFT_SOURCE_ID, order: 10 },
],
});- Use one source for ordinary “render these detections” flows.
- Use multiple sources when the app needs separate ownership, separate writes, separate retention, separate presentation, or source-aware interaction.
- Keep source IDs stable and app-owned.
- Use
requiredForPlayback: falsefor optional overlays that should never stall media playback. - Do not combine
detections.sourceswith single-source inputs such asframes,source, orappendable.