The final result of a task is visible. The way it was produced is not.
One developer says they complete tasks in one shot. Another replies that their tasks are inherently harder, cannot be completed in one shot, and the agent merely adds briefing, waiting, and corrections—making the work longer. Neither claim can be tested: people solve different tasks, in different environments, with different amounts of manual finishing. There is no shared benchmark set.
Even the subscription price explains nothing. Someone on a $20 plan may have built a highly efficient pipeline—or may simply write half the code by hand, consuming expensive human time. Someone on a $200 plan may have achieved high autonomy—or may spend the quota on personal projects and a cascade of ten checks that does not improve the result.