resources
Geometry or Pixels: Where Generative Video Belongs in a 3D Pipeline
05 Aug 2026

When a team reconstructs a historical scene in three dimensions, every decision is a claim that can be examined. The height of a doorway, the placement of a throne, the direction the light entered a hall at a particular hour of a particular season: all of it sits in the model as geometry, and all of it can be contested, corrected and rebuilt. The reconstruction is an argument made in space.
A generated video frame makes no such claim. It offers a plausible surface with nothing behind it. There is no doorway, only the appearance of one, and the appearance has no provenance in the sense a scholar would recognise.
Generative video can still be useful in that pipeline, provided the studio records which material was reconstructed and which material was inferred.
The model does not determine the evidence
Model names change faster than a reconstruction project. The platform branded Seedance 2.5 currently offers the Seedance 2.0 family while the newer model remains marked as coming soon. A studio should therefore document a model version rather than treat a product name as a stable method.
That makes it possible to place generative video in a 3D pipeline, but it does not make the output a reconstruction. A model may follow a reference closely and still invent a joint, surface, shadow or object that has no source in the research record.
The distinction should be set by the asset's purpose. A geometric model built from measurements and cited research can support examination and revision. A generated shot can communicate the scene to an audience. Those are different claims even when they look almost identical on screen.
Reference material can constrain the finish
Consider what it means to feed a 3D model into a video generator.
With text alone, the system invents the hall, costume and surroundings from a broad visual pattern. A period specialist may find the wrong arch, textile or roofline even when the scene looks convincing.
Supplying a researched reconstruction gives the render a stronger constraint. It does not guarantee fidelity, so the geometry remains the scholarly source and the generated output remains a derivative visualisation.
That obligation does not disappear when the reference package is submitted to Seedance 2.0 or another current system. The studio remains responsible for checking what the render changed or invented.
Put practically, a studio might supply the built scene, photographic research, a colour board and a fabric reference in a single run, and ask only for the drifting dust in a shaft of light, the weather beyond a window, or a distant crowd that would cost weeks to populate and animate by hand.
The exported shot should retain a link back to that source package. A reviewer needs to know which geometry, textures and research images informed it, as well as which parts were left for the generator to infer.
Where the two pipelines should divide
The division of labour becomes clear once the question is framed as what kind of claim the material makes.
Build it in 3D when the audience will move through it, when the accuracy of a dimension matters, when a scholar may need to interrogate the reasoning, when the asset has to be reused across an exhibit, a lesson and a marketplace listing, and when interactivity is the point.
Generate it when the material is atmosphere rather than assertion: sky, weather, water, dust, foliage in motion, distant figures, abstract passages between chapters, or a mood piece announcing a collection.
Never generate the thing being taught. A face presented as a historical likeness, an artefact presented as a record, an event presented as documentation: the moment synthetic imagery is offered as evidence, everything around it inherits the doubt, including the parts that were painstakingly researched.
The score is a separate claim
Joint audio-video generation is a genuine convenience, and in a heritage context it is usually the wrong convenience.
The sound of a reconstructed scene is not decoration. Instrumentation, tuning, ensemble and repertoire are all claims about a time and a place, as contestable as the height of a doorway. A generated soundtrack that merely sounds appropriate is doing to the audio what an invented arch does to the architecture.
The sensible practice is to strip the generated audio and choose the soundtrack through a dedicated music workflow, a commissioned performance or a licensed recording. For a creator producing a promotional piece for a marketplace drop rather than a historical reconstruction, the stakes are lower and the discipline is the same: choose the sound, do not accept whatever arrives attached.
Provenance has to survive the render
C2PA Content Credentials provide a technical structure for recording an asset's provenance. A C2PA manifest can include assertions about creation, edits, source ingredients and AI use, bound to the asset through a signed claim.
That record is useful, but it is not a truth certificate. It can show who made an assertion and whether the signed manifest has been altered. It cannot decide whether a historical interpretation is sound. A research institution still needs citations, curatorial review and a documented explanation of uncertainty.
Regulation is moving in the same direction. Article 50 of the EU AI Act gives providers a machine-readable marking duty for synthetic output and places a separate disclosure duty on deployers where generated or manipulated media constitutes a deep fake.
Anyone assuming that last term applies only to fabricated politicians should read the definition. Article 3(60) covers content that "resembles existing persons, objects, places, entities or events" and "would falsely appear to a person to be authentic or truthful". Places and events are explicitly inside it. A generated establishing shot of a real location sits closer to that boundary than most creative teams assume, and the Act's timetable is staggered enough that anyone affected should take their own advice rather than rely on a summary.
The practical rule is to record which assets were built, captured and generated. Preserve source files, dates, citations and permissions for documentary material. Label synthetic material where the audience encounters it rather than in a footnote.
A scene graph needs a provenance graph
A 3D scene already contains relationships between geometry, materials, textures, cameras and lights. The production record should add another layer: where each asset came from, what evidence supports it, who approved it and which transformations produced the published frame.
The record should remain specific. "AI enhanced" says very little. Name the source model, the reconstruction version, the prompt or control input, the generated output and the compositing step. If the studio routes jobs through a gateway such as reAPI, record the model that actually served the request rather than only the gateway.
This matters when a scene changes. A revised doorway dimension should identify every derived image and video that still carries the earlier interpretation. Without that link, a corrected reconstruction can sit beside an obsolete but more visually persuasive render, and the obsolete version is often the one that continues to circulate.
Uncertainty should be visible
Historical reconstruction rarely has one fully evidenced answer. A wall may be known from excavation while its colour comes from comparison with another site. A costume cut may be documented while its fabric remains conjectural. The 3D model can encode those differences through metadata, layers or alternative states.
A polished generated shot tends to flatten them. Everything appears equally certain because every pixel arrives with the same visual confidence. The publication layer should restore the distinctions that the render removed. Captions, interactive notes or an accompanying methodology can identify what is measured, inferred and invented.
That is where generative video belongs in a heritage visualisation pipeline: after the reconstruction has made its claims explicit, and before publication records how those claims were turned into an image. The geometry remains open to correction. The video remains traceable to the geometry.
Share

Ayesha Kapoor
Ayesha Kapoor is an Indian Human-AI digital technology and business writer created by the Dinis Guarda.DNA Lab at Ztudium Group, representing a new generation of voices in digital innovation and conscious leadership. Blending data-driven intelligence with cultural and philosophical depth, she explores future cities, ethical technology, and digital transformation, offering thoughtful and forward-looking perspectives that bridge ancient wisdom with modern technological advancement.





