The virtual stage
Why Narratomi places cast on five named marks and shoots them with four camera presets instead of giving you a 3D placement tool.
Open any scene and the stage is already there, whether or not you have chosen a set. It is a shallow band running left to right in front of the camera, with five positions on it. You put cast members on those positions by name. That is the whole spatial model.
Marks, not coordinates
The five marks are fixed world-space offsets, identical in every environment:
| Mark | Position on the stage line |
|---|---|
far-left | 4 metres left of center |
left | 2 metres left of center |
center | the middle of the line |
right | 2 metres right of center |
far-right | 4 metres right of center |
All five sit on the floor (y 0) at the same depth (z 0). A cast member with no mark is off-stage: not drawn, though they can still speak, and an off-stage speaker still gets a portrait. Just past the edge of the wide shot on either side are the wings, which is where a character stands when a move takes them offstage and where they enter from.
In the staging form the marks appear as five segments in stage order, abbreviated FL, L, C, R, FR, plus an off segment. The abbreviations are generated from the mark names, so the names in the bundle and the buttons in the panel can never drift apart.
Presets, not a camera rig
There are four camera presets: wide, two-shot, close-up, and over-shoulder. Each one is a shot intent, not a fixed lens position. The stage solves the actual framing per shot from who is placed and who is talking:
wideframes the stage the header set, full body: every header mark, plus any mark amoveadds. The frame holds still while a character crosses it. The establishing shot.two-shotframes the subject and their nearest scene partner, waist-up.close-upframes the subject alone, chest-up.over-shoulderputs the camera behind the partner on the subject-to-partner line, swung downstage so both read three-quarters on rather than in profile. With nobody else placed it falls back to a waist-up single.
The subject is the speaker when the speaker is on a mark, and otherwise the most center-stage character. That is why a close-up follows the conversation rather than staring at one spot. A camera direction can name a target to pin the subject to one member instead; see Framing one character.
A close-up here is really a medium shot. The dialogue box owns the bottom third of the frame, so the camera never aims above the chin line, which pins how tight any shot can get. Nothing you author changes that limit.
Framing also reacts to the viewport. The math fits the placed cast into the frustum for the current aspect ratio, so a narrow pane pulls back instead of cropping heads and name tags.
Why it works this way
The rule is short: marks and cameras live on a virtual stage, never inside environments. An environment, built-in or uploaded, is set dressing arranged around that stage. Environments carry no marks of their own, which is what makes importing one a plain upload with no placement editor behind it. The stage even hides an imported environment's own ground slab, because the floor belongs to the stage, not the set.
Three things follow.
Camera math never varies per environment. Swap the set under a scene and every shot still works, because nothing about the shot was tied to the old set's geometry.
The vocabulary is a closed list. Five marks, four camera presets, spelled the same way in the panel, in the bundle, and in the co-writer's output. An AI that only has to emit left cannot emit a broken transform matrix.
You never do 3D work to write a scene. Placement is a radio group, not a viewport you drag things around in.
The cost, stated plainly
Cast cannot be staged deep inside a set. Nobody stands on the balcony, nobody leans behind the bar, nobody stands ten metres upstage. Every scene plays on the same shallow band, visual-novel style. Depth variety is not something you can buy back with clever authoring: if a scene needs a character somewhere the marks cannot reach, put that shot in a cutaway instead, which covers the stage with an image or a video clip.
That trade is deliberate. Per-environment mark overrides are the sanctioned way to add depth later, layered on top of the default marks. Nothing in the format today lets an environment move a mark.
What the player gets
The staging header travels in the compiled bundle exactly as you filled it in. When a playthrough crosses into a scene, the runtime emits one scene-entry op carrying that header, and the stage applies it: the set loads, placed cast appear on their marks, the opening camera preset takes over, the ambient bed starts. Everything after that is stage directions in the script.
Next: staging a scene.