The difference between a prototype that flew fine all summer and an aircraft you can put in front of customers and regulators is not hours — it's structure. A flight test program converts flying into evidence: claims about the envelope, the failsafes and the reliability that someone else can audit. Here's the machinery.
The unit of work: the test card
Every productive test flight flies a card written before the crew leaves the office:
CARD 23 — Hover endurance @ MTOW, 15 kt gusts
Config: Rev C airframe, FW v1.14.2 (param set #47), payload dummy 2.0 kg
Entry: Battery ≥ 95%, wind 10–18 kt, temp 5–25 °C
Points: 1) HVR 5 m, 2 min — vibe check 2) HVR 30 m to 25% batt
Criteria: Vibe < limits; motor temps < 70 °C; sag model ±5%
Abort: Any EKF warning, temp > 80 °C, wind > 20 kt
Data: Full log rate; ESC telemetry; ambient wx recorded
Cards force pass/fail thinking, make the abort decision before adrenaline is involved, and stack into an evidence index your SORA or technical file can cite by number.
Program structure: envelope expansion
- Phase 0 — before flying: bench propulsion runs (thrust stand data), vibration survey, RF coexistence checks (integration checklist), simulation of the flight plan and every failsafe in SITL, weight-and-CG measurement against the budget.
- Phase 1 — first flights: tethered or low hover, minimum crew exposure, one variable at a time: controllability, trim, vibration, thermal trends.
- Phase 2 — envelope expansion: methodically widen mass (empty → MTOW), speed, wind, temperature, altitude, manoeuvre aggressiveness. One axis per card; corners last (MTOW + hot day + max speed is a destination, not a starting point). VTOLs add the transition corridor as its own campaign.
- Phase 3 — systems and failure injection: deliberate link loss, GNSS masking, motor-out on capable airframes, battery failsafe thresholds, geofence behaviour — at altitude, over ground that forgives.
- Phase 4 — mission and durability: the customer mission profile end-to-end, repeatedly; maintenance intervals validated; statistics accumulating (see below).
Instrumentation: if it isn't logged, it didn't happen
Modern flight stacks log superbly — the discipline is coverage and custody: full-rate logs on every flight (storage is cheaper than repeating a test), ESC telemetry and battery data included, weather actually recorded (not "breezy"), configuration captured (firmware hash, parameter set, mass) so each log is reproducible, and a post-flight review habit — ten minutes per flight scanning vibe, temps, link margins, EKF innovations catches degradation while it's still a trend, not an incident.
Safety architecture of the program itself
- Roles: even in a three-person startup, separate pilot-in-command from test director on expansion flights — the person pushing the envelope shouldn't be the one flying it.
- Site and briefing: controlled ground area appropriate to the phase, crew briefed on aborts and emergency actions, spectators managed. Your test ops should already look like your eventual authorised ops — it's rehearsal.
- Anomaly rule: any unexplained behaviour grounds the aircraft until the log explains it. "It did a weird thing but flew fine after" is how fleets learn lessons expensively.
From flights to evidence
The program's product is a growing, indexed body of claims → data: envelope limits with the cards that cleared them, failsafe demonstrations with logs, reliability statistics (flight hours, failures by subsystem, MTBF trends), maintenance findings feeding the manual. This is precisely the material BVLOS applications, class-mark files and enterprise-customer due diligence ask for — teams that structure it from flight one simply hand over a binder; teams that didn't spend a quarter reconstructing it from memory and mislabelled logs.
Flight test is where engineering optimism meets measured reality, and the program only works if reality wins every argument. Celebrate the test that finds the problem — it's the cheapest form that problem will ever take.
Frequently asked questions
How many flight hours does a new drone design need before customer use?
As a working benchmark: 20–50 hours of structured testing to clear a conservative operating envelope for supervised commercial use, and hundreds of hours of accumulated fleet time before reliability claims mean anything. What counts is coverage of the envelope and conditions, not raw hours.
What is a flight test card?
A one-page plan for a single test flight: objective, configuration, entry conditions, step-by-step test points, pass/fail criteria, abort criteria and data to record. Cards turn flying into experiments and log files into evidence.
Should failsafes be tested in flight?
Yes — deliberately and progressively. Link-loss, GNSS-degradation, low-battery and motor-out behaviour must be demonstrated, first in simulation, then at safe altitude over safe ground. A failsafe that has never fired in a test has never been tested; the first real activation shouldn't be over a customer site.