Greenlight and many conversion pipelines write their ScreenJSON to a bucket. The free screenjson-db-importer reads straight from it, with no download step.
1. Configure
Save as importer.yaml:
data_dir: ./.screenjson-importer
storage:
driver: mongo
url: mongodb://127.0.0.1:27017
database: screenplays
layout: elements # screenplays, scenes, elements, characters, analysis
blob:
driver: s3
region: us-east-1
bucket: converted-scripts
access_key: ${AWS_ACCESS_KEY_ID}
secret_key: ${AWS_SECRET_ACCESS_KEY}
For MinIO, use driver: minio and add endpoint: http://127.0.0.1:9000.
For Azure Blob Storage, use driver: azure, with the account name as
access_key and the account key as secret_key.
2. Import
screenjson-db-importer --config importer.yaml --workers 16 \
--manifest s3-import.jsonl s3://converted-scripts/2026/
Every .json object under the prefix is read, validated against the
ScreenJSON schema, and written to MongoDB. Run the same command tomorrow and
only new objects are imported; the manifest remembers the rest. The importer
never overwrites: an object whose script is already stored with different
content is reported as a failure, so edit live scripts through
screenjson-server instead.
3. Look at the result
With the default elements layout, every line is its own document: the
ScreenJSON element under node, its plain text under text, and an sj
envelope saying which script and scene it belongs to:
mongosh screenplays --eval '
db.elements.find({ "node.type": "dialogue" }, { node: 1, "sj.doc": 1 }).limit(3)
'
The scenes, characters and screenplays collections hold the rest. Use
layout: scenes for one document per scene, or layout: whole for one per
script.
Overriding settings for one run
Any setting can be given on the command line, which wins over the file:
screenjson-db-importer --config importer.yaml \
--set storage.database=screenplays_staging s3://converted-scripts/2026/
Next
- Store embeddings in a vector database
- Run screenjson-server on the same database for an API and MCP over the library.