Build the test set from real edge cases before you build the system, and the argument about quality ends.
Most teams evaluate at the end, against a sample somebody assembled that afternoon. We build the suite first, from the actual awkward cases the desk already knows about: the scanned fax, the three-way split invoice, the claim filed in the wrong language.
After that, every change runs against it. Quality stops being a matter of opinion and becomes a number on a dashboard that either moved or did not.