Text Extraction

Text Extraction

Generator provides the ability to extract the text content from your learning materials. This feature allows you to retrieve the parsed text from any piece of content that has been imported into Generator, making it useful for analysis, processing, or integration with external systems.

Overview

Generator’s text extraction feature captures all text content presented in your learning materials. This includes:

  • Course structure and organization: Section titles, chapter headers, and other organizational text
  • Slide and page content: All written content from slides, pages, and instructional material
  • Interactive elements: Quiz questions, answer choices, and feedback text
  • Embedded media assets: Transcribed text from MP3 and MP4 files, as well as parsed text from embedded PDF documents

The extracted text is organized by location within your content, making it easy to identify where specific information appears. Generator makes every effort to preserve the order of the extracted text so that it matches the order the material is presented to the learner.

Text Extraction via the API

Text extraction is performed through the GET /api/v1/content/{content_id}/versions/{version}/text endpoint.

The text extraction endpoint supports multiple output formats, each designed for different use cases:

  • TEXT: Returns the entire course’s text as plain text
  • MARKDOWN: Returns the entire course’s text as plain text, formatted in markdown
  • VTT: Returns the transcription in WebVTT format (only available for content that required transcription, such as MP3 and MP4 files)
  • TREE: Returns a more in depth representation of the course’s structure via nested json objects
  • NODES: Returns a more in depth representation of the course’s structure via a flat list of objects with parent and child ids
  • JSON: Returns the text broken up by location, including estimated token sizes
    • This format is primarily provided for legacy compatibility reasons. We generally recommend that new integrations with Generator should favor one of the other output formats, as they contain more complete information about the structure of your content.

The following information is required with each request:

  • content_id: path parameter that specifies the content identifier
  • version: path parameter that specifies the version of the content
  • tenant_id: header that specifies the tenant identifier

The following information is optional:

  • format: query parameter that specifies the output format. Accepts text, json, or vtt. Default value is text.
  • include_interactions: query parameter that specifies whether to include interaction data in the response. Default value is false. See the Interactions section below for more information.

Response Formats

TEXT Format

When using the text format, Generator returns the entire course’s text as a single plain text string. The response body contains the full text of the course as plain text, with a Content-Type of text/plain.

By default, interactions are excluded from the text. When include_interactions=true is specified, the interaction text (such as questions and answer choices) will be included in the course text at the location where the interaction appears.

This format is useful when you need the complete text content in a simple, readable format without any additional structure or metadata.

Example Response:

Welcome to this course on advanced topics...
In this chapter, we will explore...
Now that we understand the basics...

MARKDOWN Format

When using the markdown format, Generator returns the entire course’s text as a single string, with section headers denoted in markdown. The response has a Content-Type header of text/markdown.

This is similar to the text response, but text is sectioned off by the location name as a markdown header. The title of the course appears as an H1 (#) header at the top, and subsequent location names appear as H2 (##) headers after. Interactions will appear as json where they would appear in the course.

Example Response:

# Advanced Topics Course
## Introduction
Welcome to this course on advanced topics...
## Chapter 1: Getting Started
In this chapter, we will explore...
## Chapter 2: Advanced Concepts
Now that we understand the basics...

VTT Format

When using the vtt format, Generator returns the transcription in WebVTT (Web Video Text Tracks) format. The response has a Content-Type header of text/vtt.

Important:

  • The VTT format is only available for content that required transcription during the import process, such as direct MP3 audio files and MP4 video files. If you attempt to retrieve VTT format for content that does not have a transcription file (such as packaged content), the endpoint will return a 400 error with the message “Content does not have a VTT file”.
  • The VTT format does not support interactions. If you specify include_interactions=true with the VTT format, the endpoint will return a 400 error with the message “VTT format does not support interactions”.

This format is useful when you need to work with time-stamped transcriptions, create captions, or integrate with video players that support WebVTT.

Example Response:

WEBVTT

00:00:00.000 --> 00:00:05.500
Welcome to this audio course on advanced topics.

00:00:05.500 --> 00:00:12.300
In this first section, we will explore the fundamentals.

00:00:12.300 --> 00:00:20.100
Let's begin by understanding the core concepts.

TREE Format

When using the tree format, Generator returns a more in depth representation of the course’s structure as nested json objects. The response has a Content-Type header of application/json.

Rather than combining all text components into a common location, Generator presents each paragraph, header, and caption as its own text object. These text elements are nested within larger containers that reflect how the text is organized within the course. This structure is useful for anyone trying get a deeper understanding of how a slide is built, or to break down text into smaller chunks.

The TREE response contains a single root object. Each node includes:

  • id: a unique UUID for the node
  • node_type: the kind of node (content, section, text, interaction, and others)
  • source_type: how the node was obtained (typically course_data)
  • title, body, or description: the node’s localized text, when applicable
  • children: an ordered list of nested child nodes

Example Response:

{
  "id": "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d",
  "node_metadata": {},
  "token_estimation": 0,
  "node_type": "content",
  "source_type": "course_data",
  "title": {
    "locales": [
      {
        "text": "Advanced Topics Course"
      }
    ]
  },
  "children": [
    {
      "id": "3f9e2a14-6b70-4c8d-a1e5-9d0c7b4f2e18",
      "node_metadata": {},
      "token_estimation": 0,
      "node_type": "section",
      "source_type": "course_data",
      "title": {
        "locales": [
          {
            "text": "Introduction"
          }
        ]
      },
      "description": null,
      "children": [
        {
          "id": "8a1d5c32-4e90-41b7-bf26-0c3d9e7a5f14",
          "node_metadata": {},
          "token_estimation": 0,
          "node_type": "text",
          "source_type": "course_data",
          "body": {
            "locales": [
              {
                "text": "Welcome to this course on advanced topics..."
              }
            ]
          },
          "children": []
        }
      ]
    },
    {
      "id": "2b6c8e40-1d53-4a9f-8e27-5f1c0a3d7b92",
      "node_metadata": {},
      "token_estimation": 0,
      "node_type": "section",
      "source_type": "course_data",
      "title": {
        "locales": [
          {
            "text": "Chapter 1: Getting Started"
          }
        ]
      },
      "description": null,
      "children": [
        {
          "id": "9e4f1a70-2c86-4d3b-a058-1b7e6c9d4f20",
          "node_metadata": {},
          "token_estimation": 0,
          "node_type": "text",
          "source_type": "course_data",
          "body": {
            "locales": [
              {
                "text": "In this chapter, we will explore..."
              }
            ]
          },
          "children": []
        },
        {
          "id": "5d0a3c18-7f42-4e91-b6c5-8a2d1e0f3b74",
          "node_metadata": {},
          "token_estimation": 0,
          "interaction": {
            "type": "multiple_choice",
            "question": {
              "locales": [
                {
                  "text": "What is the capital of France?"
                }
              ]
            },
            "choices": [
              {
                "choice_text": {
                  "locales": [
                    {
                      "text": "London"
                    }
                  ]
                },
                "is_correct": false,
                "category": null
              },
              {
                "choice_text": {
                  "locales": [
                    {
                      "text": "Paris"
                    }
                  ]
                },
                "is_correct": true,
                "category": null
              },
              {
                "choice_text": {
                  "locales": [
                    {
                      "text": "Berlin"
                    }
                  ]
                },
                "is_correct": false,
                "category": null
              }
            ],
            "range": null,
            "feedback": []
          },
          "location": "Chapter 1: Getting Started",
          "node_type": "interaction",
          "source_type": "course_data",
          "children": []
        }
      ]
    },
    {
      "id": "1c8b5e29-3a64-4f07-9d12-6e0c4a8b7f35",
      "node_metadata": {},
      "token_estimation": 0,
      "node_type": "section",
      "source_type": "course_data",
      "title": {
        "locales": [
          {
            "text": "Chapter 2: Advanced Concepts"
          }
        ]
      },
      "description": null,
      "children": [
        {
          "id": "4f7d2b91-0e58-4c13-a6e9-2d8f1c5a0b47",
          "node_metadata": {},
          "token_estimation": 0,
          "node_type": "text",
          "source_type": "course_data",
          "body": {
            "locales": [
              {
                "text": "Now that we understand the basics..."
              }
            ]
          },
          "children": []
        }
      ]
    }
  ]
}

NODES Format

When using the nodes format, Generator returns a more in depth representation of the course’s structure as a list of json nodes. The response has a Content-Type header of application/json.

These nodes contain the same schemas that appear in the tree format. However, the elements returned in the nodes format are presented in a flat list, with lists of ids defining parents and children. This may be easier to programmatically navigate than the tree format, depending on how you’d like to process the content structure.

Each node includes:

  • id: a unique UUID for the node
  • parents: an array of parent node UUIDs (empty for the root content node)
  • children: an array of child node UUIDs
  • node_type: the kind of node (content, section, text, interaction, and others)
  • source_type: how the node was obtained (typically course_data)
  • title, body, or description: the node’s localized text, when applicable

Example Response:

[
  {
    "id": "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d",
    "parents": [],
    "children": [
      "3f9e2a14-6b70-4c8d-a1e5-9d0c7b4f2e18",
      "2b6c8e40-1d53-4a9f-8e27-5f1c0a3d7b92",
      "1c8b5e29-3a64-4f07-9d12-6e0c4a8b7f35"
    ],
    "node_metadata": {},
    "token_estimation": 0,
    "title": {
      "locales": [
        {
          "text": "Advanced Topics Course"
        }
      ]
    },
    "node_type": "content",
    "source_type": "course_data"
  },
  {
    "id": "3f9e2a14-6b70-4c8d-a1e5-9d0c7b4f2e18",
    "parents": [
      "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d"
    ],
    "children": [
      "8a1d5c32-4e90-41b7-bf26-0c3d9e7a5f14"
    ],
    "node_metadata": {},
    "token_estimation": 0,
    "title": {
      "locales": [
        {
          "text": "Introduction"
        }
      ]
    },
    "description": null,
    "node_type": "section",
    "source_type": "course_data"
  },
  {
    "id": "8a1d5c32-4e90-41b7-bf26-0c3d9e7a5f14",
    "parents": [
      "3f9e2a14-6b70-4c8d-a1e5-9d0c7b4f2e18"
    ],
    "children": [],
    "node_metadata": {},
    "token_estimation": 0,
    "body": {
      "locales": [
        {
          "text": "Welcome to this course on advanced topics..."
        }
      ]
    },
    "node_type": "text",
    "source_type": "course_data"
  },
  {
    "id": "2b6c8e40-1d53-4a9f-8e27-5f1c0a3d7b92",
    "parents": [
      "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d"
    ],
    "children": [
      "9e4f1a70-2c86-4d3b-a058-1b7e6c9d4f20",
      "5d0a3c18-7f42-4e91-b6c5-8a2d1e0f3b74"
    ],
    "node_metadata": {},
    "token_estimation": 0,
    "title": {
      "locales": [
        {
          "text": "Chapter 1: Getting Started"
        }
      ]
    },
    "description": null,
    "node_type": "section",
    "source_type": "course_data"
  },
  {
    "id": "9e4f1a70-2c86-4d3b-a058-1b7e6c9d4f20",
    "parents": [
      "2b6c8e40-1d53-4a9f-8e27-5f1c0a3d7b92"
    ],
    "children": [],
    "node_metadata": {},
    "token_estimation": 0,
    "body": {
      "locales": [
        {
          "text": "In this chapter, we will explore..."
        }
      ]
    },
    "node_type": "text",
    "source_type": "course_data"
  },
  {
    "id": "5d0a3c18-7f42-4e91-b6c5-8a2d1e0f3b74",
    "parents": [
      "2b6c8e40-1d53-4a9f-8e27-5f1c0a3d7b92"
    ],
    "children": [],
    "node_metadata": {},
    "token_estimation": 0,
    "interaction": {
      "type": "multiple_choice",
      "question": {
        "locales": [
          {
            "text": "What is the capital of France?"
          }
        ]
      },
      "choices": [
        {
          "choice_text": {
            "locales": [
              {
                "text": "London"
              }
            ]
          },
          "is_correct": false,
          "category": null
        },
        {
          "choice_text": {
            "locales": [
              {
                "text": "Paris"
              }
            ]
          },
          "is_correct": true,
          "category": null
        },
        {
          "choice_text": {
            "locales": [
              {
                "text": "Berlin"
              }
            ]
          },
          "is_correct": false,
          "category": null
        }
      ],
      "range": null,
      "feedback": []
    },
    "location": "Chapter 1: Getting Started",
    "node_type": "interaction",
    "source_type": "course_data"
  },
  {
    "id": "1c8b5e29-3a64-4f07-9d12-6e0c4a8b7f35",
    "parents": [
      "7c4a8d09-0e3b-4f2a-9c1d-2e6f8a0b1c2d"
    ],
    "children": [
      "4f7d2b91-0e58-4c13-a6e9-2d8f1c5a0b47"
    ],
    "node_metadata": {},
    "token_estimation": 0,
    "title": {
      "locales": [
        {
          "text": "Chapter 2: Advanced Concepts"
        }
      ]
    },
    "description": null,
    "node_type": "section",
    "source_type": "course_data"
  },
  {
    "id": "4f7d2b91-0e58-4c13-a6e9-2d8f1c5a0b47",
    "parents": [
      "1c8b5e29-3a64-4f07-9d12-6e0c4a8b7f35"
    ],
    "children": [],
    "node_metadata": {},
    "token_estimation": 0,
    "body": {
      "locales": [
        {
          "text": "Now that we understand the basics..."
        }
      ]
    },
    "node_type": "text",
    "source_type": "course_data"
  }
]

JSON Format

When using the json format, Generator returns the text broken up by location, with each location including its text content and an estimated token count. The response has a Content-Type header of application/json.

The JSON response follows this schema:

  • locations: a dictionary where each key is a location string and each value is a TextLocationResponse object containing:
    • text: the text content for that location
    • token_estimation: an estimated token count for the text at that location
    • interactions: (optional) an array of interaction objects. This field is only present when include_interactions=true. See the Interactions section for details on the interaction structure.
  • token_estimation: the total estimated token count for all locations

By default, interactions are excluded from the response. When include_interactions=true is specified, interactions are included in a separate interactions field for each location.

This format is primarily provided for legacy compatibility reasons. We generally recommend that new integrations with Generator should favor one of the other output formats, as they contain more complete information about the structure of your content.

Example Response (without interactions):

{
  "locations": {
    "Introduction": {
      "text": "Welcome to this course on advanced topics...",
      "token_estimation": 45
    },
    "Chapter 1: Getting Started": {
      "text": "In this chapter, we will explore...",
      "token_estimation": 120
    },
    "Chapter 2: Advanced Concepts": {
      "text": "Now that we understand the basics...",
      "token_estimation": 200
    }
  },
  "token_estimation": 365
}

Example Response (with interactions):

{
  "locations": {
    "Introduction": {
      "text": "Welcome to this course on advanced topics...",
      "token_estimation": 45,
      "interactions": []
    },
    "Chapter 1: Getting Started": {
      "text": "In this chapter, we will explore...",
      "token_estimation": 120,
      "interactions": [
        {
          "type": "multiple_choice",
          "question": {
            "locales": [
              {
                "text": "What is the capital of France?"
              }
            ]
          },
          "choices": [
            {
              "choice_text": {
                "locales": [
                  {
                    "text": "London"
                  }
                ]
              },
              "is_correct": false
            },
            {
              "choice_text": {
                "locales": [
                  {
                    "text": "Paris"
                  }
                ]
              },
              "is_correct": true
            },
            {
              "choice_text": {
                "locales": [
                  {
                    "text": "Berlin"
                  }
                ]
              },
              "is_correct": false
            }
          ],
          "feedback": []
        }
      ]
    },
    "Chapter 2: Advanced Concepts": {
      "text": "Now that we understand the basics...",
      "token_estimation": 200,
      "interactions": []
    }
  },
  "token_estimation": 365
}

Content Structure

There are many types of nodes which can be used to represent the content structure in the tree and nodes formats. To learn more about the various node types, see Content Structure .

Interactions

Many e-learning authoring tools include interactive elements such as quizzes, questions, and assessments within their content. Generator can extract and include these interactions in the text extraction output when the include_interactions parameter is set to true. See here for more information about interactions.

Limitations

To ensure the highest accuracy, the current version of Generator focuses on core text extraction. The following elements are outside the current scope of our extraction process at this time:

  • Images: Visual content, including charts, diagrams, infographics, and photographs, cannot be extracted as text. Only alt text or image descriptions will be included if provided.
  • External links: URLs and external web content they reference are not fetched or included.
  • Embedded external videos: Video content hosted on YouTube and Vimeo are not extracted.
  • Formatted styling: Text extraction preserves content but not visual formatting such as colors, fonts, emphasis, or layout positioning.