> For the complete documentation index, see [llms.txt](https://scade.gitbook.io/scade-knowledge-base/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://scade.gitbook.io/scade-knowledge-base/flow-examples/building-a-flow-video-transcription-and-summarization.md).

# Building a flow: video transcription and summarization

Life is too short for watching long videos? Let's build a workflow to transcribe and summarize any Youtube video.

1. Go to the [Flow](https://app.scade.pro/flow) and click the "Create" button. <br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FtJbMYrqDrHESxN85QF8m%2Fimage.png?alt=media&amp;token=dd457d47-6af8-47a3-93e9-459810e302b4" alt=""><figcaption></figcaption></figure>

   <br>
2. It's a good practice to store all your input data in a single node. Start by typing `user input` in the search bar of the left panel. Locate **User-Defined Input Form** and drag it to the workspace.<br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F7jOnlScowMt8CQT2iH2z%2FScreenshot%202024-02-16%20at%2020.58.12.png?alt=media&amp;token=e42f6b14-202f-4e9d-ae99-10e62d087198" alt=""><figcaption></figcaption></figure>

   <br>
3. This node has no pre-configured fields, so we should set them ourselves. Dive into the node’s settings and select **Configure Fields**. <br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F6bxps0gvSKxu6WSar0BM%2FScreenshot%202024-02-16%20at%2021.03.01.png?alt=media&amp;token=b6e38872-9a7f-45e0-b270-a1139eba21c3" alt=""><figcaption></figcaption></figure>

   <br>
4. Add a field to store the link to the video. Let's name this field `video_url`. We don't have to change the field type or anything else. Click **Save**.\
   &#x20;

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FBNZ0LflULBvanLiYOAeY%2FScreenshot%202024-02-16%20at%2021.20.51.png?alt=media&amp;token=ab156654-c51f-46b7-a624-c2fe50b96eb2" alt=""><figcaption></figcaption></figure>

   <br>
5. Paste the Youtube link to the node, hit the **Execute** button and then click **Save and execute** to run this node and generate the output. We need it to move further.\
   \
   A video that is used as an example in this flow: <https://www.youtube.com/watch?v=JVatgo0TJIw><br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F9bXdCSaaAcDCwCu0OHqM%2FScreenshot%202024-02-16%20at%2021.29.50.png?alt=media&amp;token=a89fbcca-4364-42c0-9bf6-aacb4540a6e4" alt=""><figcaption></figcaption></figure>

   \
   \
   Here's the output<br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F9hNrWdGMTkJVuz3lJMAW%2FScreenshot%202024-02-16%20at%2021.36.33.png?alt=media&amp;token=0c8618a4-81d0-4469-999e-955de2e3a1ab" alt=""><figcaption></figcaption></figure>

   <br>
6. Start typing in the search bar what you want to do next: `video transcription`. We've found two models, let's give one of them a try. Drag **whisperx-video-transcribe** to your workspace<br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FCOpy61HD2BGyeKCOtVe3%2FScreenshot%202024-02-16%20at%2021.40.39.png?alt=media&amp;token=131df539-11cf-44d3-bf91-875d915a1cec" alt=""><figcaption></figcaption></figure>

   <br>
7. Сonnect the **video\_url** output of your **User-Defined Input Form** node to the **Url** input of your **whisperx-video-transcribe** node. <br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2Fs3btQJixz3Ypau4bkr1i%2FScreenshot%202024-02-16%20at%2021.45.01.png?alt=media&amp;token=e4ef1a6b-3a58-4e9a-b760-c6364b84dcbb" alt=""><figcaption></figcaption></figure>

   <br>
8. Execute the **whisperx-video-transcribe** node. It will take a while to transcribe a video; the longer the video is, the more time is needed. \ <br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FqOGh3WqI2IdEVNKewIwo%2FScreenshot%202024-02-16%20at%2021.49.38.png?alt=media&amp;token=5841a143-2707-44fa-97d2-bb622d6bbc19" alt=""><figcaption></figcaption></figure>

   <br>
9. We have a transcription, let's summarize it. Start typing `chatgpt processor` in the left panel and drag a **ChatGPT Processor** node to your workspace.\ <br>

   <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FoZ6nLLsWtE4UAU7rVcpx%2FScreenshot%202024-02-16%20at%2021.56.50.png?alt=media&amp;token=b074cf9d-f11f-400e-91db-90f05dd04158" alt=""><figcaption></figcaption></figure>

   <br>
10. Go to settings of the **ChatGPT Processor** node. \
    \
    You can choose ChatGPT versions with the **Model** dropdown. Let's choose gpt-4 because the transcript could be too long for older versions<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FKxPqIL0k7LKCFhhEVQhI%2FScreenshot%202024-02-16%20at%2022.02.45.png?alt=media&amp;token=c485475e-9e8b-4fb2-a66f-c83a6b80466a" alt=""><figcaption></figcaption></figure>

    <br>
11. Now it's time to prepare ChatGPT for receiving instructions. \
    \
    In the node settings, go to **Messages** section, click on a pencil and a new message<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FaloNb7H2ZYpgsUgM2ATv%2FScreenshot%202024-02-16%20at%2022.05.56.png?alt=media&amp;token=0c184e31-2988-4ebb-b145-36dc1c792282" alt=""><figcaption></figcaption></figure>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2Fo5tRIRynDSIr71D7xTfx%2FScreenshot%202024-02-16%20at%2022.06.34.png?alt=media&amp;token=cf7263fc-ad32-457b-bf4d-1c699f612874" alt=""><figcaption></figcaption></figure>

    \
    \
    Change the type of the message to `System` as recommended by OpenAI to emphasize that it's a high-level instruction and to guide the model's behaviour.<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F04gh5ygld2pCuWK88xO3%2FScreenshot%202024-02-16%20at%2022.07.15.png?alt=media&amp;token=f8ed333b-d742-4333-9d35-d4390c7f8913" alt=""><figcaption></figcaption></figure>

    <br>
12. In the **Message** field write what you need **ChatGPT** to do: something like `summarize a following video transcript` <br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FYik6SRk7V0BKBveGL94M%2FScreenshot%202024-02-16%20at%2022.12.23.png?alt=media&amp;token=92610691-7419-4b5f-af6f-c2f6bbc460a0" alt=""><figcaption></figcaption></figure>

    <br>
13. Add another message (there's no need to change the message type this time). \
    \
    Here's a tricky part: due to the syntax of the **ChatGPT Processor** node we will need to use **Expression editor**. \
    \
    Click on the **#** symbol in the lower **Message** field.<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2Fyaa6QHmCFAUbtyTMOXux%2FScreenshot%202024-02-16%20at%2022.21.05.png?alt=media&amp;token=2e6f5deb-1ef0-4882-ba5c-c61dd8c6edc9" alt=""><figcaption></figcaption></figure>

    \
    \
    You'll get the **Expression editor**. There is a list of the nodes on the left. Click on the **whisperx-video-transcribe** to see its output.<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2F2zjqH4sKK1ICx4THyQXx%2FScreenshot%202024-02-16%20at%2022.33.51.png?alt=media&amp;token=79e0bd73-b389-498e-8b05-56f31f849d93" alt=""><figcaption></figcaption></figure>

    \
    \
    Drag the **success** output to the expression field<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FjZzvhYkqXzg6WXcHToN5%2FScreenshot%202024-02-16%20at%2022.34.17.png?alt=media&amp;token=e95499dc-57e4-4576-a48a-5d7b8fb2a26e" alt=""><figcaption></figcaption></figure>

    \
    \
    The result should look like this<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FWiIoA9W5o8YRAjFoi406%2FScreenshot%202024-02-16%20at%2022.35.10.png?alt=media&amp;token=83d5d9cf-852a-4097-b897-1bb4b48fbd7d" alt=""><figcaption></figcaption></figure>

    \
    \
    Click the **Save** button in the **Expression editor** and don't forget to click top-right **Save** as well before leaving the settings.<br>
14. Connect the output of the **whisperx-video-transcribe** to the input of the **ChatGPT Processor** node and run the latter by clicking the **Execute** button on the node. \
    \
    An important note: don't click the **Start** button on the top panel yet! If you do this, all nodes will start over, and we don't want this because video transcription is a quite time-consuming task.<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FcEHzxgX3L9QHTVg9r6BI%2FScreenshot%202024-02-16%20at%2022.41.31.png?alt=media&amp;token=baacee60-36dd-43ab-ba58-56f29268d29f" alt=""><figcaption></figcaption></figure>

    <br>
15. When the **ChatGPT** finishes its work, you'll see that the text in this node is way shorter than the initial transcript. You can read it right away or copy it — just hover over the text to see the respective icon. <br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FGalgN6zU7PAQE8UH12kf%2FScreenshot%202024-02-16%20at%2023.11.29.png?alt=media&amp;token=280b271c-ddd6-47aa-b225-c1e65dfa7f54" alt=""><figcaption></figcaption></figure>

    \
    \
    Of course, you can copy the full transcript as well.<br>
16. Like storing all input data in a single node, collecting all output data in a single node is a good practice as well. For example, it will simplify matters greatly if you are going to launch the workflow via API in the future.\
    \
    Let's grab another **User-Defined Input Form** and add a couple of fields. Name them, for example, `full_transcript` and `summary`. Refer to paragraphs 2, 3 and 4 of this guide if needed.\ <br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FfmMX7cvQ9lKrOjvasu5v%2FScreenshot%202024-02-16%20at%2023.22.33.png?alt=media&amp;token=b3256157-e012-4284-bc42-be9e004a88ba" alt=""><figcaption></figcaption></figure>

    <br>
17. Connect the output of the **whisperx-video-transcribe** node to the **Full\_transcript** input and the **Success** output of the **ChatGPT Processor** to the **Summary** input respectively. Run the final node by clicking the **Execute** button\ <br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FtgskWGOPDrtq0aNWv4zj%2FScreenshot%202024-02-16%20at%2023.26.04.png?alt=media&amp;token=704f615b-7534-49c0-909c-8abc02535298" alt=""><figcaption></figcaption></figure>

    <br>
18. Now both a full transcript and a summary are stored in one node.<br>

    <figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2FibZh0BrROJDGhnzTNgpq%2FScreenshot%202024-02-16%20at%2023.39.49.png?alt=media&amp;token=8fd79628-dba4-410f-bc91-4bdbddb3e01d" alt=""><figcaption></figcaption></figure>

    <br>

And finally: once you have the workflow built, there's no need to run nodes one by one anymore. To transcribe and summarize another video, simply change the link in the **Video\_url** field of the **User-Defined Input Form** and use the **Start** button on the top bar this time.<br>

<figure><img src="https://2793209830-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FNCrOWKpBlXYnNB9SFIzb%2Fuploads%2Fq1xE3tfaR2tqcWl8xgyO%2FScreenshot%202024-02-16%20at%2023.42.57.png?alt=media&amp;token=3a75af22-7b9f-4e1a-9bd9-a15f6c8844a4" alt=""><figcaption></figcaption></figure>
