<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AI Research Notes</title>
    <link>https://wangsheng1991.github.io/content-engine/</link>
    <description>Short, reproducible write-ups of AI research and models, each one backed by a runnable artifact.</description>
    <language>en</language>
    <lastBuildDate>2026-09-22T12:42:07Z</lastBuildDate>
    <atom:link href="https://wangsheng1991.github.io/content-engine/feed.xml" rel="self" type="application/rss+xml"/>

    <item>
      <title>Qwen-Image-2.1 on public prompts: the 4/7 was protocol, not capability</title>
      <link>https://wangsheng1991.github.io/content-engine/en/blog/qwen-image-2-1-bench/</link>
      <guid isPermaLink="true">https://wangsheng1991.github.io/content-engine/en/blog/qwen-image-2-1-bench/</guid>
      <pubDate>2026-09-22</pubDate>
      <description>Qwen-Image-2.1, SenseNova-U1.5 and a 4-step INT4 FLUX.2-klein, on one machine and one GPU, run against the verbatim text of GenEval / DPG-Bench / LongText-Bench, same seed: Qwen passed 4 of 7 compositional prompts where both baselines went 7/7. But the recommended pipeline is prompt rewriter + 2048² + 40 steps, and round one used none of them. Rewriting the prompts into enhancer-style descriptions — same resolution, same seed — fixed both object-dropping failures on the spot. These are prompts the model does not eat, not pictures it cannot draw. A second, self-authored set of 26 vertical prompts (9 domains, 156 images) answers &quot;which model should take this job&quot;: no single model wins everywhere. SenseNova takes renders, portraits, UI mockups and comic panels; Qwen owns the Chinese poster headline and native transparent RGBA.</description>

      <enclosure url="https://wangsheng1991.github.io/content-engine/assets/qwen-image-2-1-bench/og.en.png" type="image/png" length="38943"/>

    </item>

    <item>
      <title>Image-to-video in three parameters: how a 5-second clip costs twice what it should</title>
      <link>https://wangsheng1991.github.io/content-engine/en/blog/wan-i2v-first-frame/</link>
      <guid isPermaLink="true">https://wangsheng1991.github.io/content-engine/en/blog/wan-i2v-first-frame/</guid>
      <pubDate>2026-09-22</pubDate>
      <description>Wan 2.6 video-from-image (wan2.6-i2v-flash) is the easiest rung: one first frame plus one motion prompt, and five seconds later you have an MP4. The API is asynchronous only — task IDs and result URLs both expire after 24 hours. The expensive part is two other parameters: resolution is a pixel budget rather than a quality setting (and defaults to the dearer 1080P), and audio defaults to on, so a 5-second 720P clip is billed at the with-sound rate — exactly double, ¥1.50 instead of ¥0.75. Our three June clips were billed that way: every file carries an AAC track.</description>

      <enclosure url="https://wangsheng1991.github.io/content-engine/assets/wan-i2v-first-frame/og.en.png" type="image/png" length="42930"/>

    </item>

  </channel>
</rss>
