← 深度专栏/产品观察
产品观察

The Cycling Pelican: What a Quirky Test Reveals About AI Economics

How do you accurately measure the evolutionary leap between two generations of artificial intelligence? You could run complex mathematical benchmarks, analyze...

潜
作者
潜龙编辑部
关注 AI 与社会议题
发布于
2026/10/5
READ
长读
The Cycling Pelican: What a Quirky Test Reveals About AI Economics
illustration · QianLong editorial

How do you accurately measure the evolutionary leap between two generations of artificial intelligence? You could run complex mathematical benchmarks, analyze millions of data points, or parse through dense academic papers. Or, if you want a more practical demonstration, you could simply ask the AI to draw a pelican riding a bicycle using raw SVG code.

Recently, developer Simon Willison used this exact, delightfully quirky prompt to stress-test OpenAI’s new GPT-6 Astra model, comparing its output against the previous generation’s GPT-5.6 family, which includes models like Sol, Terra, and Luna. By asking the AI to write the scalable vector graphics (SVG) code to render this scene at various reasoning levels, he created a visual comparison grid that reveals a lot more than just an AI's artistic limitations.

The most immediate takeaway from the experiment is the sheer jump in spatial reasoning and output quality. When generating SVGs, the AI isn't "painting" pixels like Midjourney; it is writing mathematical coordinates and shapes. The older GPT-5.6 Sol model, even when pushed to its maximum reasoning capabilities, produced what looked like a jumble of abstract geometric shapes. In stark contrast, GPT-6 Astra’s absolute lowest reasoning setting generated a highly recognizable, superior image. When cranked to its "max" setting, Astra's output was exceptionally good.

However, the experiment also showed that AI still struggles with basic physical logic. Unless Astra was set to its absolute highest reasoning tier, it repeatedly failed to place the pelican’s legs on opposite sides of the bicycle frame—a classic example of how large language models can excel at complex code while stumbling on simple real-world physics.

But the most compelling revelation from the "cycling pelican" test isn't about art; it's about economics. On paper, Astra looks like an expensive upgrade. It costs roughly $10 per million input tokens and $50 per million output tokens, which is double the price of the older Sol model ($5 input / $30 output).

Yet, the actual cost per task tells a completely different story. Because Astra possesses superior reasoning capabilities, it requires significantly fewer tokens to achieve a successful result. In Willison's test, generating a high-quality pelican with Astra on a low setting cost just 9.55 cents. Spending a similar 10 cents on the older, "cheaper" models yielded a vastly inferior result. Furthermore, a look at the input token counts showed Astra and Luna both using exactly 16 tokens for the prompt, compared to 26 for Sol and Terra, hinting at a potential shared architectural lineage between specific models.

This dynamic flips the traditional tech upgrade narrative on its head. As AI models become more sophisticated, their sticker price per token may increase, but their operational efficiency can actually make them cheaper to use in practice. For businesses and developers, the cycling pelican serves as a perfect reminder: in the world of generative AI, working smarter ultimately costs less.

Key Points

  • SVG generation serves as a rigorous stress test for an AI's spatial reasoning and coding capabilities.
  • GPT-6 Astra's lowest setting outperformed the maximum capability of the older GPT-5.6 Sol model.
  • Higher token pricing for advanced models is often offset by their efficiency, resulting in lower per-task costs.
  • Despite advanced coding skills, AI models still occasionally struggle with basic real-world physics, like spatial positioning.

Why It Matters

This highlights a counterintuitive reality in AI adoption: upgrading to a model with a higher base price can actually reduce operational costs due to increased token efficiency.


Sources:

潛
本文完
潜龙编辑部 · 2026/10/5
潜龙 QianLong · 中文 AI 内容与工具平台