Content Structure
The tree and nodes text formats present a much higher fidelity to how text and other elements within courses are structured. Generator will represent text as various nodes in a graph. See
Text Extraction
for how to request these formats.
Shared fields on every node include id (a UUID), node_type, source_type, token_estimation, node_metadata, and children. Text-bearing fields (title, body, description) use a TextElement.
NOTE: Generator’s ability to extract data from courses is constantly improving. With future updates we expect to add more schemas and better leverage the schemas we currently have.
The node types are as follows:
TextElement
While not technically its own node, TextElements are a common pattern that appear within nodes when documenting text within the course. TextElements are a pattern intended to support courses that contain multiple language options for the same text.
A TextElement contains a locales array. Each entry is one variant of the text. Entries are not unique by locale: several strings may share a null or identical locale.
Each entry has a text string. When the language is known, the entry also includes source_declared_locale (the locale declared in the course), detected_locale (the locale Generator identified), and/or text_direction (ltr or rtl). These fields are omitted when they are null.
{
"locales": [
{
"source_declared_locale": "en-US",
"detected_locale": "en-US",
"text_direction": "ltr",
"text": "Welcome to this course."
},
{
"source_declared_locale": "fr",
"detected_locale": "fr",
"text_direction": "ltr",
"text": "Bienvenue dans ce cours."
},
{
"text": "Bienvenido a este curso."
}
]
}
Content
The root node for a content version. Its title is the course title. All other nodes sit under this root.
{
"id": "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d",
"node_type": "content",
"source_type": "course_data",
"title": {
"locales": [
{
"text": "Advanced Topics Course"
}
]
},
"children": []
}
Section
A generic container for modules, slides, pages, or other groupings. Sections can contain text, interactions, files, or other sections.
{
"id": "3f9e2a14-6b70-4c8d-a1e5-9d0c7b4f2e18",
"node_type": "section",
"source_type": "course_data",
"title": {
"locales": [
{
"text": "Introduction"
}
]
},
"description": null,
"children": []
}
Text
A single block of learner-facing text, such as a paragraph or caption. Text nodes do not have children.
{
"id": "8a1d5c32-4e90-41b7-bf26-0c3d9e7a5f14",
"node_type": "text",
"source_type": "course_data",
"body": {
"locales": [
{
"text": "Welcome to this course on advanced topics..."
}
]
},
"children": []
}
Interaction
A quiz or other interactive element. See
Interactions Extraction
for the interaction payload.
{
"id": "5d0a3c18-7f42-4e91-b6c5-8a2d1e0f3b74",
"node_type": "interaction",
"source_type": "course_data",
"location": "Chapter 1: Getting Started",
"interaction": {
"type": "multiple_choice",
"question": {
"locales": [
{
"text": "What is the capital of France?"
}
]
},
"choices": [
{
"choice_text": {
"locales": [
{
"text": "Paris"
}
]
},
"is_correct": true,
"category": null
}
],
"range": null,
"feedback": []
},
"children": []
}
File
An embedded media or document file. file_type is one of video, audio, image, pdf, vtt, or other. file_path is a path inside the course package, or a remote URL.
{
"id": "6e2f9a11-4c80-4b3d-9e17-0a5d8c2b1f40",
"node_type": "file",
"source_type": "course_data",
"file_type": "video",
"file_path": "video/intro.mp4",
"media_resolved": true,
"children": []
}
VTT file
A WebVTT caption file packaged with the course. Its children are vtt_cue nodes.
{
"id": "0b7c4e25-8d91-4a16-bf03-2e9a1c6d5f38",
"node_type": "vtt_file",
"source_type": "course_data",
"file_type": "vtt",
"file_path": "captions/intro.vtt",
"media_resolved": true,
"children": []
}
Transcription
A transcription of audio or video. source_type is AWS_TRANSCRIBE when Generator produced the transcript, or course_data when the captions came from the course. Its children are vtt_cue nodes.
{
"id": "9c3a1d70-2e54-4f18-a6b9-1d0e7c4f8a25",
"node_type": "transcription",
"source_type": "AWS_TRANSCRIBE",
"file": "video/intro.mp4",
"media_resolved": true,
"children": []
}
VTT Cue
A timed caption within a transcription or vtt_file. start_time and end_time use WebVTT timestamp format. source_type matches the parent (AWS_TRANSCRIBE or course_data).
{
"id": "4a8e2c16-7f30-4d95-b1a2-6c0d9e3f5b18",
"node_type": "vtt_cue",
"source_type": "course_data",
"identifier": null,
"location": "00:00:00.000",
"start_time": "00:00:00.000",
"end_time": "00:00:05.500",
"body": {
"locales": [
{
"text": "Welcome to this audio course on advanced topics."
}
]
},
"children": []
}