A fair AI video comparison is not two attractive clips placed side by side. It is a small experiment in which another student can understand the question, repeat the procedure and see how the conclusion was reached.
MiniMax H3 and Wan 3.0 do not expose identical controls, so “same settings” cannot mean forcing every menu value to match. The defensible approach is to hold the task constant, match the controls that overlap, record the differences that cannot be removed, and score several outputs without declaring a winner from one lucky generation.
The method below produces a dataset, scoring sheet, limitations section and seminar presentation—not a promotional ranking.
1. Write One Narrow Research Question
Start with a question that the experiment can actually answer:
Under matched six-second, 16:9 text-to-video tasks, how do MiniMax H3 and Wan 3.0 differ in prompt adherence, subject continuity, motion continuity, physical plausibility, audio alignment and visible artifacts?
This compares text-to-video only, not image-to-video or reference-video performance. Naming the dimensions before generating also prevents criteria from changing after the results appear.
Use a neutral hypothesis: “The models will show different strengths across the six criteria.” A benchmark does not need a predicted champion.
2. Define the Test Matrix Before Generating
Use three task types, three repeats per task and two prompt conditions. With two models, this creates 36 clips:
3 tasks × 3 repeats × 2 prompt conditions × 2 models = 36 outputs
Three repeats expose run-to-run variation but cannot support strong statistical claims. If budget is tight, reduce task types rather than repeats.
|
Task ID |
What the prompt tests |
Observable event |
|
T1: handoff |
object interaction and contact |
one person passes a red notebook to another |
|
T2: bicycle |
lateral motion and background continuity |
cyclist crosses a junction as the camera pans |
|
T3: café line |
dialogue timing, small gestures and ambience |
seated speaker says one short sentence, lifts a cup, room tone continues |
Avoid celebrities, logos and copyrighted characters. In Condition A, use the same prompt, duration and ratio. Do not enable prompt enhancement for only one model.
In Condition B, keep the event, setting, duration and shot boundary unchanged, but follow each model’s current prompt guidance. This checks model-specific instruction structure without changing the task.
3. Match What Can Be Matched
Six seconds and 16:9 sit inside both models’ documented text-to-video ranges. Resolution is less tidy: the available tiers are not named identically. Record the requested tier and the delivered pixel dimensions for every file. If the dimensions differ, do not score sharpness as if resolution had been controlled.
Use this control sheet before the first run:
|
Variable |
Rule |
|
Mode |
text-to-video only |
|
Duration |
6 seconds requested |
|
Aspect ratio |
16:9 |
|
Prompt condition |
A: identical; B: model-guided structure |
|
Audio |
enabled for both; score separately from visuals |
|
Repeats |
3 per task, condition and model |
|
Generation order |
alternate models to reduce time-of-day bias |
|
Evaluation |
random file codes; model names hidden from raters |
|
Editing |
none beyond uniform playback preparation |
Record any different duration, frame rate or file size as an outcome. Keep raw files untouched; make separately labelled playback copies if the presentation requires them.
4. Pilot the Prompts Without Scoring Them
Run one unscored pilot for each task. The purpose is to catch broken wording: perhaps the handoff has no clear giver, the camera instruction contradicts the staging, or the spoken line is too long for six seconds.
Enter the three draft tasks in ClipDance and make a pilot sheet that records the selected model, duration, ratio, prompt version and any attached input; revise only instructions that make the event ambiguous, then freeze version 1.0 before the measured run.
Pilots may be discarded because their role was defined in advance. Once measurement starts, every completed output belongs in the dataset, including failures.
5. Collect a Reproducible Run Log
For the measured runs, submit the frozen matrix through reAPI and capture the returned task identifier, model identifier, submission time, completion time, prompt condition, requested controls, terminal status and reported usage in a CSV row. Keep the original output filename tied to that task identifier.
A compact CSV header is enough:
run_id,task_type,repeat,condition,model_id,prompt_version,duration_requested,
ratio_requested,resolution_requested,task_id,submitted_at,completed_at,status,
duration_observed,width,height,fps,has_audio,reported_usage,file_name
Do not put API keys, signed links or student login details in the sheet. Save the full prompts in a separate text file with stable labels such as T1-A-v1.0 and T1-B-H3-v1.0.
Alternate model order across repeats. This cannot eliminate service changes, but it avoids collecting the two model sets on widely separated dates.
6. Score Blind, Then Reveal the Labels
Rename viewing copies with random codes. Ask at least two classmates who did not write the prompts to score independently. Show the rubric before playback and let each rater replay every clip the same number of times.
Use a 0–4 scale:
- 0: task failed or the required element is absent.
- 1: major errors dominate.
- 2: requirement is partly met, with obvious errors.
- 3: requirement is met, with minor errors.
- 4: requirement is met cleanly throughout the clip.
|
Run code |
Adherence |
Subject continuity |
Motion continuity |
Physics |
Audio alignment |
Artifacts |
Notes |
|
Q17 |
|||||||
|
M04 |
Define each category in one sentence. “Subject continuity” covers clothing, body and key objects. “Motion continuity” means actions progress without unexplained resets. “Audio alignment” covers timing and relevance, not taste. Notes should cite evidence, such as “notebook changes colour at 00:03.”
Calculate each criterion’s median and range by model and task. An overall average can hide task-specific differences. Keep latency and reported usage in separate factual tables, outside the subjective score.
7. State the Limitations Before the Results
A strong limitations section should include the following points:
- The sample has only three prompts and three repeats, so results are descriptive.
- Text-to-video findings do not apply automatically to image or reference modes.
- Resolution tiers and delivered dimensions may not match exactly.
- Model-guided prompts in Condition B require researcher judgement.
- Human ratings contain subjectivity even with a rubric and blinding.
- Hosted model revisions may change during or after collection.
- Latency includes routing and queue conditions, not model computation alone.
- A college budget may be too small to estimate rare failure rates.
Record the access date and finish the runs in the shortest practical window. If either model changes during collection, note the date and treat the dataset as two periods rather than concealing the interruption.
8. Seminar Slide Outline
- Problem: why one side-by-side clip is weak evidence.
- Research question: the exact scope and six criteria.
- Model context: MiniMax H3 and Wan 3.0 controls used in this study.
- Design: 3 tasks, 3 repeats, 2 prompt conditions and 2 models.
- Rubric: 0–4 definitions and blind-rating procedure.
- Results: criterion medians, ranges and completion data.
- Failure examples: two short clips with time-coded observations.
- Limitations: mismatched tiers, small sample and service revisions.
- Findings: task-specific results, with no universal winner.
- Reproduction packet: prompts, run log, rubric and access date.
Frequently Asked Questions
Is using the same prompt for both models automatically fair?
No. It is one useful condition, but identical wording may favour one model’s preferred instruction style. Pair it with a second condition that follows each model’s current guidance while preserving the same observable event.
How many generations are enough for a college project?
Three repeats per cell are a practical minimum for showing that outputs vary, not a basis for broad statistical certainty. State the sample size prominently and report ranges rather than hiding variation behind one mean.
Should cost and speed decide the benchmark winner?
Only if the research question says they should. Report completion time and usage separately. A student can discuss trade-offs, but combining money, latency, audio and visual quality into one unexplained total makes the result hard to interpret.
What Earns the Marks
The value of this project lies in its records, not a dramatic verdict. Freeze the task matrix, preserve every measured run, score blind and show the limits beside the results. Another student should be able to follow the folder, prompts and CSV from research question to conclusion. That is what makes a MiniMax H3 versus Wan 3.0 comparison a benchmark rather than a highlight reel.
