This AI video creation workflow combines five specialized tools: ChatGPT for planning and script revision, Midjourney for visual development, D-ID for an avatar-led video, vidIQ for YouTube packaging and research, and Audacity for authorized audio recording and editing.
The original 2023 video calls it a four-tool workflow but also uses Audacity, so the accurate count is five. Several interfaces and model versions have changed since the recording. This updated guide preserves the demonstrated sequence while distinguishing archived controls from the current products.
Watch the complete historical workflow above, or jump to the contextual timestamps for image creation, avatar animation, YouTube optimization, script drafting, and audio recording.
Disclosure: the original creator URLs from the video description—including the tracked vidIQ destination—are preserved exactly. AI Tools Arena may receive a referral benefit. Product features and plan limits can change.
The Five-Tool AI Video Workflow
| Tool | Current role | Output passed to the next stage |
|---|---|---|
| ChatGPT | Audience research, outline, script drafting, and revision | Verified scene-by-scene script |
| Midjourney | Character or scene concept image | Approved, rights-checked image |
| D-ID | Presenter or avatar-led video generation | Generated and reviewed video clip |
| Audacity | Record or edit audio you are authorized to use | Clean narration or final audio track |
| vidIQ | YouTube ideas, keyword research, titles, descriptions, and packaging support | Human-approved metadata and publishing plan |
The tools do not replace the production decisions between them. Every handoff needs a review: factual accuracy before narration, visual rights before animation, consent before avatar use, audio rights before editing, and audience fit before publishing.
Step 1: Plan and Write the Script With ChatGPT
Start with the viewer, not the software. Define who the video is for, the problem it solves, the evidence it will show, and the action viewers should take. Provide ChatGPT with approved source material and explicitly prohibit invented facts.
A useful script prompt is:
Create a [duration] YouTube script for [audience] about [topic]. The viewer should be able to [outcome]. Use only the source material below and flag missing information instead of guessing. Organize the script into scenes with narration, on-screen text, visual direction, and source notes. Keep sentences natural for speech and do not invent prices, results, quotations, or product features. Source material: [paste approved sources].
The archived video demonstrates ChatGPT script drafting at 02:59.
Verify before generating video
- Check names, dates, statistics, features, prices, and instructions against primary sources.
- Remove repetitive introductions and unsupported superlatives.
- Read the script aloud and shorten difficult sentences.
- Mark pronunciation for acronyms and brand names.
- Separate spoken narration from on-screen text.
- Identify every visual the viewer needs as evidence.
Do not let a fluent draft create false confidence. A synthetic presenter can make an inaccurate statement appear authoritative.
Step 2: Create the Visual in Midjourney
The original tutorial uses Midjourney V4, the beta upscaler, --q 5, and a 16:9 ratio. Those settings belong to a legacy workflow. Midjourney’s current official version documentation lists V8.2 as the default at this update.
Watch the original Midjourney V4 image-generation step at 00:23.
Use the old prompt as an archive, not a current settings recipe
most beautiful long black hair, pale skin, emerald green eyes, Greek attire, flawless skin, show her full head, show her full face, show her half body, dramatic lighting, minimalist, wearing tank top, realistic, cinematic, photography, natural skin, delicate, detailed, photo realistic, beauty photography, 4k, sci-fi –v 4 –ar 16:9 –upbeta –q 5
The subject and composition can still inspire a new prompt, but remove unsupported legacy parameters and use current documentation for version, quality, upscaling, references, and personalization. The current default can change again, so do not hard-code a model version into a permanent production template without a review date.
Current Midjourney visual checklist
- Specify subject, environment, framing, lighting, medium, and aspect ratio.
- Generate several variations and change one major instruction at a time.
- Inspect eyes, hands, clothing, jewelry, text, logos, and background geometry.
- Check whether a reference image or style reference is authorized.
- Use the current upscaler documented for the selected model.
- Save the prompt, model, settings, date, and approved output.
For a detailed current workflow, see our Midjourney AI tutorial.
Step 3: Animate the Presenter With D-ID
The video demonstrates uploading a character image to D-ID, adding a script, selecting a language and voice, and generating the clip. D-ID’s current official feature page still describes pre-scripted avatar video creation, plus stock, personal, and custom avatars, AI voices, voice cloning, and localization. D-ID now also promotes real-time visual agents, which are a different use case from a rendered presenter video.
Watch the D-ID image-to-presenter workflow at 00:53.
D-ID production steps
- Choose an image you own or are licensed to animate.
- Confirm that the person depicted has authorized the intended avatar use.
- Paste the verified narration or upload authorized audio.
- Select the language and voice.
- Generate a short test before the full script.
- Check lip sync, facial movement, expression, pronunciation, and image artifacts.
- Export using the current account’s available settings.
Do not present a synthetic avatar as a real recording of a person. Obtain consent, restrict account access, and label synthetic media when context, platform policy, or law requires it.
Step 4: Record or Edit Authorized Audio in Audacity
Audacity is a free, open-source desktop audio editor and recorder for Windows, macOS, and Linux. The official download page should be used for the current stable version.
The 2023 video records computer playback with Audacity at 03:18. That technique is technically possible on supported systems, but it is not automatically permitted. Recording a web service’s playback can reduce quality and may violate licensing or platform terms.
Safer audio workflow
- Prefer the platform’s authorized export when it provides one.
- Record your own microphone or media you have permission to capture.
- Keep an unedited source track.
- Remove mistakes and excessive silence.
- Apply noise reduction conservatively to avoid metallic artifacts.
- Normalize or adjust levels without clipping.
- Export a lossless master and the delivery format required by the editor.
Audacity’s official support documents describe Windows WASAPI loopback for desktop audio. macOS does not provide built-in desktop-audio capture in the same way and may require an authorized routing setup. Follow the current operating-system guide rather than copying a Windows setting to every device.
Step 5: Package the Video With vidIQ
The original workflow asks vidIQ to create a title, description, tags, and hashtags. Current vidIQ pages continue to offer keyword research, title and description generators, scripts, thumbnail ideas, optimization recommendations, and other YouTube tools.
Watch the archived vidIQ optimization step at 01:54.
Important correction to the old tutorial
The video says to ask vidIQ to “check Ahrefs” for keywords. Do not assume vidIQ has access to an Ahrefs account or live Ahrefs data. Use vidIQ’s own current keyword and optimization tools, and validate recommendations against YouTube search results, your channel analytics, and the actual video.
YouTube packaging checklist
- Title: state the result or test accurately without promising more than the video shows.
- Thumbnail: make the subject understandable at small size and do not fake the result.
- Description: answer the topic quickly, summarize the workflow, and preserve creator or affiliate links exactly.
- Chapters: use verified timestamps from the final uploaded cut.
- Tags: add a limited set of relevant spelling, entity, and topic variants; tags do not compensate for weak content.
- Disclosure: identify affiliate links and synthetic or altered media when required.
Use AI suggestions as candidates. The creator remains responsible for choosing packaging that matches the real video and audience.
Recommended Production Order
- Define the audience, objective, and source of truth.
- Draft the script with ChatGPT.
- Fact-check and approve the script.
- Create and review the Midjourney visual.
- Generate a short D-ID avatar test.
- Edit only authorized audio in Audacity.
- Assemble the scenes and complete captions.
- Watch the exported video on desktop and mobile.
- Use vidIQ for research and metadata candidates.
- Approve the title, thumbnail, description, links, and chapters manually.
AI Video Quality-Control Matrix
| Stage | Risk | Required check |
|---|---|---|
| ChatGPT script | Invented or outdated claims | Verify every fact against primary sources |
| Midjourney image | Distorted anatomy, logos, or objects | Inspect full resolution and usage rights |
| D-ID avatar | Misleading identity or poor lip sync | Confirm consent, disclosure, and visual quality |
| Audacity audio | Unauthorized capture or overprocessing | Confirm rights and listen for artifacts |
| vidIQ metadata | Clickbait or unsupported keyword claims | Match every promise to the final video |
| Final export | Timing, caption, or layout failure | Review on target devices before upload |
How to Modernize the 2023 Workflow
| Archived step | Current approach |
|---|---|
| Midjourney V4 and beta upscaler | Use the current default and compatible upscaler from official documentation |
| One long portrait prompt | Test subject, framing, reference, and style decisions separately |
| D-ID as a simple talking-photo tool | Choose between pre-scripted avatar video and current real-time agent products |
| Record web playback by default | Prefer authorized exports; use Audacity for owned or permitted audio |
| Ask vidIQ to check Ahrefs | Use vidIQ’s own data and verify with YouTube and channel analytics |
| Publish generated metadata | Human-review every promise, link, keyword, and disclosure |
Frequently Asked Questions
Does this workflow use four or five tools?
Five. The original title says four, but the actual workflow uses ChatGPT, Midjourney, D-ID, vidIQ, and Audacity.
Can I still use the Midjourney V4 prompt?
You can study it as a historical example, but V4 and its legacy parameters are not the current default workflow. Rebuild the prompt for the current model and verify compatible settings.
Do I need Audacity if D-ID exports a video?
Not always. If D-ID provides the authorized final audio and video you need, avoid unnecessary capture. Audacity remains useful for recording your own narration or editing permitted audio.
Will vidIQ guarantee YouTube ranking?
No. It can support research and packaging, but performance depends on viewer response, topic demand, competition, satisfaction, and the quality and accuracy of the video.
Final Verdict
The five-tool workflow still offers a useful production model because each tool has a distinct job. Its weak point is the handoff between tools: an unverified script becomes a confident synthetic voice, a flawed image becomes an animated presenter, or an AI title promises a result the video never shows.
Use current official documentation, preserve rights and consent, test each asset before the next stage, and review the final export as one complete experience. That turns an old collection of AI tricks into a repeatable, accountable content workflow.
