Archive notice: This comparison was recorded in March 2023 and tests the GPT-4 and earlier ChatGPT experience available at that time. It is preserved as a same-prompt experiment, not as a guide to the models or plans currently offered in ChatGPT.
The original test used 12 identical prompts across creative writing, website generation, planning, social content, and simple app development. That methodology is still useful: when comparing AI models, keep the prompt, context, and evaluation criteria constant so differences are easier to identify.
Watch the original GPT-4 vs. GPT-3 comparison below to inspect the outputs demonstrated in the video. No new result or winner has been invented for this archived article.
What This 2023 GPT-4 vs. GPT-3 Test Covered
The comparison was designed to answer a practical question at the time: did GPT-4 produce a more useful result than the earlier ChatGPT model when both received the same request? Rather than relying on benchmark scores alone, the video tested deliverables a creator might actually request.
| Task category | Examples in the test | Useful evaluation criteria |
|---|---|---|
| HTML and coding | CV website and mobile calculator | Valid code, responsiveness, interaction, accessibility, and whether the file runs |
| Creative writing | Romantic poems and a Bermuda Triangle horror story | Originality, coherence, tone, repetition, and instruction following |
| Marketing content | Lipstick caption and SEO company presentation | Audience fit, specificity, structure, persuasive clarity, and unsupported claims |
| Short-form ideas | Pickup lines and fictional Tinder profiles | Variety, tone, safety, usefulness, and repetition |
| Informational writing | Mental-health article | Accuracy, sourcing, safety, organization, and appropriate limitations |
| Planning | YouTube schedule and Bali itinerary | Feasibility, sequencing, assumptions, current information, and completeness |
| Video scripting | YouTube yoga script | Hook, flow, audience level, factual safety, pacing, and call to action |
The 12 Original Same-Prompt Tests
- Design a website to showcase an online curriculum vitae with sample content in one HTML file.
- Create a mobile-friendly HTML5 calculator app in one HTML file.
- Create one HTML file containing five romantic poems.
- Create one HTML file containing a horror story set in the Bermuda Triangle.
- Write an Instagram caption for a lipstick product and display it on a designed webpage.
- Develop a presentation for an SEO company and display it on a webpage.
- Generate 10 pickup lines and present them on a website.
- Create an SEO-friendly mental-health article and present it on a website.
- Create 10 humorous fictional Tinder profiles and display them on a playful website.
- Create a daily schedule for a productive YouTube creator and display it on a structured website.
- Plan a one-week Bali vacation itinerary and present it on a website.
- Create a YouTube script about practicing yoga and present it on a website.
The wording above has been cleaned for readability, but the task intent remains the same as the original 2023 experiment.
What Makes a Fair AI Model Comparison?
Use the exact same input
Keep the prompt, source material, files, and constraints identical. Even a small wording change can alter the result enough to make the comparison unreliable.
Start a clean session
Previous messages can influence later responses. Use a new conversation for each model unless conversation memory is the feature you intentionally want to test.
Define the rubric before generating
Choose measurable criteria such as requirement coverage, factual accuracy, code execution, source quality, readability, originality, and revision time. Do not decide what matters only after seeing which output looks better.
Test more than one run
Generative outputs vary. One response can be unusually strong or weak. Repeating the prompt helps separate a consistent capability difference from random variation.
Verify the finished deliverable
Run the code, open every link, check calculations, verify current facts, and read the full response. A polished presentation can hide broken behavior or unsupported claims.
Why This Comparison Is No Longer a Buying Guide
The product has changed substantially since 2023. Current ChatGPT documentation describes newer model families, different modes and tools, and a broader set of Free and paid plans. Access also depends on the plan, product surface, workspace, and current usage limits.
For today’s product, read our current ChatGPT review. For plan pricing and upgrade criteria, see Is ChatGPT free? Current plans and limits.
How to Repeat the Test With Current Models
- Choose two models that are actually available in the same product surface and account.
- Copy one original prompt into a new session for each model.
- Add the same source files and constraints.
- Generate at least three outputs per model.
- Score each output with a predefined rubric.
- For HTML tasks, save and run every file in a browser.
- For factual tasks, verify claims against current primary sources.
- Record the model name, date, settings, plan, and tool access.
- Compare not only first-draft quality but also how many revisions were required.
That final metric matters. A model that produces a slightly prettier first answer may still be less useful if it requires more correction, ignores constraints, or creates code that does not run.
Frequently Asked Questions
Is GPT-4 still the current ChatGPT model?
No. This article documents a 2023 comparison. Consult the current model selector and official ChatGPT documentation for the models available in your account today.
Can the 12 prompts still be reused?
Yes, as test inputs. Update vague requirements, add accessibility and safety checks, and verify any task involving current information. For mental health, travel, or other consequential topics, require reliable sources and human review.
Does one same-prompt test prove which model is better?
No. It shows how the models responded to one request under one set of conditions. A useful conclusion needs repeated runs, several task types, and a consistent scoring method.
Final Takeaway
The lasting value of this GPT-4 vs. GPT-3 video is its test design: use identical prompts and evaluate real deliverables. Its model-access and upgrade instructions belong to 2023 and should not guide a current purchase. Use the archived prompts to build a repeatable comparison, then judge today’s models by accuracy, requirement coverage, execution, and revision effort.
