<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://spencer.imbleau.com/blog/feed.xml" rel="self" type="application/atom+xml" /><link href="https://spencer.imbleau.com/blog/" rel="alternate" type="text/html" /><updated>2026-09-08T05:48:48+00:00</updated><id>https://spencer.imbleau.com/blog/feed.xml</id><title type="html">Spencer Imbleau’s blog</title><subtitle>Writing by Spencer Imbleau about Rust, computer graphics, open-source software, game development, AI, and engineering.</subtitle><author><name>Spencer Imbleau</name></author><entry><title type="html">AI is in its homelab era</title><link href="https://spencer.imbleau.com/blog/ai/2026/09/07/ai-is-still-in-its-homelab-era.html" rel="alternate" type="text/html" title="AI is in its homelab era" /><published>2026-09-07T00:00:00+00:00</published><updated>2026-09-07T00:00:00+00:00</updated><id>https://spencer.imbleau.com/blog/ai/2026/09/07/ai-is-still-in-its-homelab-era</id><content type="html" xml:base="https://spencer.imbleau.com/blog/ai/2026/09/07/ai-is-still-in-its-homelab-era.html"><![CDATA[<p>Gone are the days my mom would ask me for technical help with her printer; she has ChatGPT. So too are the days my brother would ask me if I could make a website for his brick-and-mortar business; he knows it’s possible to vibe code one.</p>

<p>I don’t think I need to tell you that AI is popular—it is. We are in an AI popularity bubble. But I do not think people are prepared for where it is going. AI is nowhere near saturated.</p>

<p>My brother’s requests changed from <em>“Could you build this website?”</em> to <em>“Can you vibe this website in a few minutes?”</em> Nature is healing, or nothing really changed. People still expect software to cost, be secure and hygienic, and come with support. AI has not yet markedly convinced non-technical people that they can make it themselves.</p>

<p>We have popularity without saturation.</p>

<h2 id="the-expensive-hobby-phase">The expensive hobby phase</h2>

<p>My AI usage can punch upwards of $3,000 in a month. Most of it is coding, since my craft is software engineering, while a very small amount is personal AI use. Most of my AI coding use will stay the same or increase—I believe coding everything by hand will not be generally competitive for long. However, even on my software engineering team, there are varying levels of saturation. In the graph below, my usage is roughly four times higher than the next-highest coworker. I do not attribute this to working four times as hard; I see it as evidence that even software engineers have not reached high usage.</p>

<figure class="evidence-figure">
  <a href="/blog/assets/ai-team-usage-recent.svg" target="_blank" rel="noopener noreferrer">
    <img src="/blog/assets/ai-team-usage-recent.svg" alt="Two horizontal bar charts comparing recent AI traces and AI cost for seven anonymized software engineers. Spencer leads in cost and traces." />
  </a>
  <figcaption>AI traces and attributed AI cost across my anonymized software engineering team.</figcaption>
</figure>

<p>This makes me a terrible example of normal AI use, but a useful example of its upper end.</p>

<p>I am not using personal AI tools much. I have tried OpenClaw, but I found its usefulness shallower than the public perception. Instead, I find it’s easier than ever to create exactly what I want - zero to one. I have my own daily agenda app with some modest AI capability. I have made my own messaging bots I can interact with to perform certain tasks.</p>

<p>But no tool in existence yet integrates my development workflow on a scalable platform and manages my day. Most tools I use are tools I visit, or maintain, and occasionally forget about.</p>

<p><em>And this feels familiar</em>. Five years ago, my personal infrastructure was a homelab built from five Raspberry Pis running k3s with RAID storage, redundant networking, independent power switches, and a cooling solution. It was fun. But it also cost hundreds to thousands of dollars and required enough maintenance to qualify as a small, unpaid operations team.</p>

<p>Now my homelab collects dust in my garage, and my personal projects live in my AWS account. That includes this website, my daily agenda app with a database, some business projects, a few Discord bots, and more. My monthly AWS bill is on average $0.09.</p>

<p>Innovation made that possible. Services like Amazon DSQL—a distributed, serverless relational database with zero idle cost—became generally available and gave me a reasonable home for my daily TODOs, health data, and more.</p>

<p>The homelab was useful, but maintenance-heavy and expensive. The cloud turned this into a cheap commodity. I expect AI to follow the same curve.</p>

<h2 id="cheap-enough-to-forget">Cheap enough to forget</h2>

<p>AI is prohibitively expensive when used as aggressively as I use it now. That will not last. Models will get smaller, hardware will improve, inference will become more competitive, and providers will keep finding cheaper ways to serve the same useful unit of intelligence. <a href="https://epoch.ai/data-insights/llm-inference-price-trends" target="_blank" rel="noopener noreferrer">Capability has been getting cheaper</a>, and that is likely to continue. <a href="https://www.anthropic.com/news/higher-limits-spacex" target="_blank" rel="noopener noreferrer">Anthropic raising free usage limits</a> may be another sign of that pressure reaching users.</p>

<p>My prediction is not merely that AI gets cheaper. It becomes cheap enough that ordinary people stop thinking about the meter.</p>

<p>That affects how I think about the broad AI trade in public markets. This is not investment advice. I am personally wary of betting on scarcity when the product looks destined for commoditization. Subsidized competition from China and other countries could force prices down further, although some protected markets—government and defense come to mind—may behave differently.</p>

<h2 id="saturation-looks-like-a-personal-manager">Saturation looks like a personal manager</h2>

<p>AI reaches saturation when most people use it to manage their lives, not infrequent use of a chatbot like Gemini.</p>

<p>Everyone will need something like an OpenClaw in roughly the same way everyone now needs a phone or access to the internet. It will be competitive, and it will understand the annoying surface area of a day: messages, appointments, tasks, files, forms, reminders, purchases, and the small promises we make before immediately forgetting them.</p>

<p>It has not yet arrived for me, even though the AI usage attributable to me could finance a bad habit.</p>

<p>I am, however, coding personal tools more frequently. I see this trend growing independently; “AgenticOS” tutorial series are becoming quite popular on YouTube.</p>

<figure class="evidence-figure">
  <img src="/blog/assets/agentic-os-youtube-collage.jpg" alt="Collage of YouTube thumbnails advertising Claude AgenticOS tutorials" />
  <figcaption>AgenticOS tutorials on YouTube.</figcaption>
</figure>

<p>These “OS” tutorials are really just slop dashboards, but I get the allure. Who wouldn’t want JARVIS from Iron Man? Here is the gap - doing this right requires scaling, maintenance, and hard integration work. These are fun experiments for AI enthusiasts. A useful personal manager would need a phone companion, sensible alerts, durable memory, authorization, and recovery when an agent does something stupid. Most people watching these tutorials should not need to learn how containers work before asking software to remember a dentist appointment.</p>

<p>And even worse, OpenClaw. There was a magic spark when I tried it, but that faded quickly. The interface was slop, and its security model did not earn my personal data.</p>

<p>Someone will eventually make the Apple-like quality version of OpenClaw: narrow enough to understand, polished enough to trust, and useful before the user has configured forty integrations. When that happens, the world will come to it. Until then,</p>

<h2 id="a-hostile-launch-environment">A hostile launch environment</h2>

<p>A loud part of the public is vehemently against AI.</p>

<p>People see companies exploiting the political system, building data centers that may raise local water and electricity costs, expecting a government bailout, and burning cash without a believable return. Some of those concerns are measurable. Others are projections. A credible personal AI platform will have to survive all of them.</p>

<p>Saturation will only make this launch environment more hostile. Adoption and acceptance are pulling in opposite directions: more aggregate compute means more electricity, cooling, and water demand. <a href="https://hai.stanford.edu/ai-index/2026-ai-index-report" target="_blank" rel="noopener noreferrer">Global nervousness about AI</a> is rising while excitement drops. Models and data centers will get more efficient, but cheaper intelligence will also invite us to use much more of it.</p>

<p>I think we will literally need an act of Congress: national rules deciding when data centers pay for the infrastructure they require, how scarce resources are allocated, and what costs can be passed to everyone else. AI companies want government assistance, but it needs to be politically favorable, or public opposition to AI will continue to harden.</p>

<p>The winning product cannot merely be capable. It must be cheap, boring, secure, and visibly worth the infrastructure behind it.</p>

<h2 id="five-years-probably">Five years, probably</h2>

<p>My guess is that we are at least five years away from saturation, unless another ChatGPT-sized viral moment pulls the schedule forward.</p>

<p>By around 2031, I expect a credible personal agent to be common on phones and personal devices. It should:</p>

<ul>
  <li>behave like a normal app;</li>
  <li>have understandable permissions and mistakes;</li>
  <li>cost little enough that use is not a financial decision; and</li>
  <li>become a meaningful part of the day.</li>
</ul>

<p>The final signal will not be another benchmark. It will be my brother building the software for his shop without asking me anything.</p>

<p>And if he leaks the customer database, we are only halfway there.</p>]]></content><author><name>Spencer Imbleau</name></author><category term="ai" /><summary type="html"><![CDATA[AI is popular, but not saturated. My predictions for how cost, adoption, and public backlash shape the next five years.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://spencer.imbleau.com/og-banner.png" /><media:content medium="image" url="https://spencer.imbleau.com/og-banner.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Press: ‘Compute-centric vector graphics using bevy_vello’</title><link href="https://spencer.imbleau.com/blog/rust/graphics/2024/08/01/press-bevy_vello.html" rel="alternate" type="text/html" title="Press: ‘Compute-centric vector graphics using bevy_vello’" /><published>2024-08-01T00:00:00+00:00</published><updated>2024-08-01T00:00:00+00:00</updated><id>https://spencer.imbleau.com/blog/rust/graphics/2024/08/01/press-bevy_vello</id><content type="html" xml:base="https://spencer.imbleau.com/blog/rust/graphics/2024/08/01/press-bevy_vello.html"><![CDATA[<p>I recently was invited to be a speaker at the <a href="https://www.meetup.com/bevy-game-development/">Bevy Meetup</a>, sponsored by <a href="https://rustunit.com/">Rustunit</a>. The recording is now availble, and you may find extra material below. <em>Thank you to <a href="https://rustunit.com/">https://rustunit.com/</a></em>.</p>

<iframe width="100%" height="400px" src="https://www.youtube.com/embed/VQGQhotekvY" frameborder="0" allowfullscreen=""></iframe>

<h2 id="download">Download</h2>

<ul>
  <li><a href="https://docs.google.com/presentation/d/14tBxPvYxchgvgJ7tejyfW__h3pPqoJbhLmpHAeFL0Bs">Slides (Google Slides)</a></li>
  <li><a href="/blog/assets/Compute-centric-vector-graphics-with-bevy_vello.pdf">Slides (download .pdf)</a></li>
  <li><a href="/blog/assets/Compute-centric-vector-graphics-with-bevy_vello.pptx">Slides (download .pptx)</a></li>
</ul>

<h2 id="citation">Citation</h2>

<p>Please use the following attribution, if needed:</p>

<ul>
  <li><strong>BibTeX</strong></li>
</ul>

<div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code>@misc<span class="p">{</span>imbleau-2024,
    author = <span class="p">{</span>Imbleau, Spencer<span class="p">}</span>,
    month = <span class="p">{</span>8<span class="p">}</span>,
    title = <span class="p">{{</span>Compute-centric vector graphics using bevy<span class="p">_</span>vello<span class="p">}}</span>,
    year = <span class="p">{</span>2024<span class="p">}</span>,
    url = <span class="p">{</span>https://www.youtube.com/watch?v=VQGQhotekvY<span class="p">}</span>,
<span class="p">}</span>
</code></pre></div></div>

<ul>
  <li><strong>BibLaTeX</strong></li>
</ul>

<div class="language-tex highlighter-rouge"><div class="highlight"><pre class="highlight"><code>@video<span class="p">{</span>imbleau-2024,
    author = <span class="p">{</span>given-i=SCI, given=Spencer, family=Imbleau<span class="p">}</span>,
    date = <span class="p">{</span>2024-08-01<span class="p">}</span>,
    title = <span class="p">{</span>Compute-centric vector graphics using bevy<span class="p">_</span>vello<span class="p">}</span>,
    url = <span class="p">{</span>https://www.youtube.com/watch?v=VQGQhotekvY<span class="p">}</span>,
<span class="p">}</span>
</code></pre></div></div>

<ul>
  <li><strong>APA</strong> (<em>7th Edition</em>)</li>
</ul>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Imbleau, Spencer &amp; Rustunit. (2024, August 1). Compute-centric vector graphics using bevy_vello [Video]. YouTube. https://www.youtube.com/watch?v=VQGQhotekvY
</code></pre></div></div>]]></content><author><name>Spencer Imbleau</name></author><category term="rust" /><category term="graphics" /><summary type="html"><![CDATA[Lessons from making the first modern compute-centric vector graphic game for web]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://spencer.imbleau.com/og-banner.png" /><media:content medium="image" url="https://spencer.imbleau.com/og-banner.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">High performance vector graphic video games</title><link href="https://spencer.imbleau.com/blog/rust/graphics/2023/11/20/using-vello-for-video-games.html" rel="alternate" type="text/html" title="High performance vector graphic video games" /><published>2023-11-20T00:00:00+00:00</published><updated>2023-11-20T00:00:00+00:00</updated><id>https://spencer.imbleau.com/blog/rust/graphics/2023/11/20/using-vello-for-video-games</id><content type="html" xml:base="https://spencer.imbleau.com/blog/rust/graphics/2023/11/20/using-vello-for-video-games.html"><![CDATA[<p>For a few years I’ve been trying to solve a hard problem: <em>“How can I use vector graphics as the backing image model in realtime systems?”</em></p>

<p>Yes, the image model that uses points, lines, and equations to encode image data. Of which, several common encoding formats exist, such as like SVG and PDF.</p>

<p><img src="/blog/assets/Bitmap_VS_SVG.svg" alt="svg" /></p>

<p>The internet has done this once meaningfully in the past. In the 2000s, a herculean effort to optimize the performance of interactive vector graphics brought an advent of vector games through Adobe’s Shockwave Flash player. But in the wake of SWF’s deprecation from major browsers, there wasn’t a worthy successor, and the era of web games died off. Unfortunately, efforts like HTML5 and WebGL were never able to replace what was lost.</p>

<p>Hence, I believe the first games to understand and successfully use this image model under the constraints of major browsers could benefit from content-delivery savings, resolution independence, and offer novel applications of interactivity, physics, and non-destructive transformations.</p>

<h2 id="onto-the-papercuts">Onto the papercuts</h2>

<p>Vector images are notoriously unfit for modern GPU architecture because of an inherent locality issue. In contrast to raster graphics where color and fill information is explicitly encoded, rendering a pixel in a vector image requires knowledge of the entire image. Efficiently solving the fill of every pixel in realtime poses a unique challenge. With the rising accessibility of compute kernels and low-level GPU architecture access over the past few years, especially from projects like <a href="https://wgpu.rs/">wgpu</a>, friction with general purpose GPU computing is fading. Projects like <a href="https://github.com/linebender/vello">vello</a> are pioneering the 2d vector graphics space, and furthering the experimental research inspired by projects like <a href="https://github.com/servo/pathfinder">pathfinder</a>.</p>

<p>The high-level strategy used by these new hardware-accelerated renderers is to create a <a href="https://raphlinus.github.io/rust/graphics/gpu/2020/06/12/sort-middle.html">compute-centric pipeline with a sort-middle architecture</a> to parallelize the problem space efficiently. This is a technical leap in recent years and its efficiency will be hard to challenge.</p>

<h2 id="vong">Vong</h2>

<p><img src="/blog/assets/vong.png" alt="Vong" /></p>

<p>The technical leap was implemented in <a href="https://github.com/linebender/vello">vello</a> (<em>formerly piet-gpu</em>), started by <a href="https://levien.com/">Raph Levien</a> and my colleagues in <a href="https://linebender.org">the linebender community</a>. The Vello API depends on WebGPU to provide an abstraction layer to targets, while offering more accessibility to low-level hardware like compute shaders. It’s for these reasons rust, webgpu, and vello would make a perfect combination for cross-platform vector games with one code-base.</p>

<p><img src="/blog/assets/webgpu.svg" alt="Vong" /></p>

<p>This spurred on my race to prove the lost benefits are still relevant. In a sudden case of reviving-the-horse syndrome, I worked with a few colleagues from my previous tenure at NASA to understanding the unknowns of why I’ve never seen a modern <a href="https://www.behance.net/kurzgesagt">Kurzgesagt</a> vector game. Hence, the decision to start by creating the classic game of Pong was deliberate, providing a foundational exploration of Vello’s capabilities in rendering vector graphics dynamically. It felt appropriate to call the project <em>Vong</em>, joining “Vector” and “Pong”.</p>

<p>Matching webgpu, vello, and rust’s strengths, <a href="https://bevyengine.org/">bevy</a> was the obvious choice for a cross-platform open source game engine. It also uses webgpu, and was developed open-source and code-first.</p>

<h3 id="rendering-integration">Rendering integration</h3>

<p>Standing in the way of seeing my first <a href="https://commons.wikimedia.org/wiki/File:Ghostscript_tiger_(original_background).svg">Ghostscript Tiger</a> (the <em>Hello, World!</em> of vector graphics) was an integration to render vector assets in bevy with vello. Sebastian Hamel (<a href="https://github.com/seabassjh">@seabassjh</a>) and I (<a href="https://github.com/nuzzles">@nuzzles</a>) developed that integration in the open, <a href="https://github.com/vectorgameexperts/bevy-vello"><code class="language-plaintext highlighter-rouge">bevy-vello</code></a>. Just like any other ordinary asset in bevy, such as a <em>png</em>, there are no surprises:</p>

<div class="language-rust highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">// Bevy 0.12</span>
<span class="k">fn</span> <span class="nf">setup_assets</span><span class="p">(</span>
    <span class="k">mut</span> <span class="n">commands</span><span class="p">:</span> <span class="n">Commands</span><span class="p">,</span>
    <span class="n">asset_server</span><span class="p">:</span> <span class="n">ResMut</span><span class="o">&lt;</span><span class="n">AssetServer</span><span class="o">&gt;</span>
<span class="p">)</span> <span class="p">{</span>
    <span class="c1">// Load image of egg</span>
    <span class="n">commands</span><span class="nf">.spawn</span><span class="p">(</span>
        <span class="n">VelloVectorBundle</span> <span class="p">{</span>
            <span class="n">vector</span><span class="p">:</span> <span class="n">asset_server</span><span class="nf">.load</span><span class="p">(</span><span class="s">"egg.svg"</span><span class="p">),</span>
            <span class="n">origin</span><span class="p">:</span> <span class="nn">bevy_vello</span><span class="p">::</span><span class="nn">Origin</span><span class="p">::</span><span class="n">Center</span><span class="p">,</span>
            <span class="n">transform</span><span class="p">:</span> <span class="nn">Transform</span><span class="p">::</span><span class="nf">from_scale</span><span class="p">(</span><span class="nn">Vec3</span><span class="p">::</span><span class="nf">splat</span><span class="p">(</span><span class="mf">0.1</span><span class="p">)),</span>
            <span class="o">..</span><span class="nn">Default</span><span class="p">::</span><span class="nf">default</span><span class="p">()</span>
        <span class="p">}</span>
    <span class="p">)</span>
<span class="p">}</span>
</code></pre></div></div>

<p>Architecturally, <code class="language-plaintext highlighter-rouge">bevy-vello</code> was built with two backends. The first is an <code class="language-plaintext highlighter-rouge">.svg</code> ETL library called <a href="https://github.com/vectorgameexperts/vello-svg"><code class="language-plaintext highlighter-rouge">vello-svg</code></a> which loads the SVG using the vello <code class="language-plaintext highlighter-rouge">SceneBuilder</code> API. The second is <a href="https://github.com/vectorgameexperts/vellottie"><code class="language-plaintext highlighter-rouge">vellottie</code></a>, a parser and runtime for <a href="https://airbnb.io/lottie/">Airbnb’s <code class="language-plaintext highlighter-rouge">.json</code> lottie</a> format, and a rewrite of <a href="https://github.com/linebender/velato">Chad Brokaw’s lottie runtime, “velato”</a>.</p>

<iframe width="100%" height="400px" src="https://www.youtube.com/embed/hNu5oF18j5g" frameborder="0" allowfullscreen=""></iframe>

<p>Strategically, it made sense to pursue both formats, but a painful lesson learned was that the <code class="language-plaintext highlighter-rouge">.svg</code> and <code class="language-plaintext highlighter-rouge">.json</code> lottie specifications are huge, with plenty of feature flags that aren’t supported by vello. While writing a parser wasn’t too challenging using the <code class="language-plaintext highlighter-rouge">serde</code> parsing library, dealing with animation inconsistencies and rendering artifacts remains a difficult task. Currently, vello is writing CPU shaders to help triage these artifacts. But for the future, I’m inclined to support only lottie formats, since any SVG <em>could</em> be reformed into a 0 frame-rate lottie animation. One of the biggest obstacles in commercial viability is artist tooling, so I’m keeping an eye on companies like <a href="https://lottielab.com">LottieLab</a>, <a href="https://phase.com">Phase</a>, and <a href="https://rive.app">Rive</a>.</p>

<h3 id="physics">Physics</h3>

<p>Once rendering was solved, another thorny issue with pong was physics: collision detection between arbitrarily curved shapes is <em>really hard</em>. While some algorithms exist, none today are obviously good.</p>

<p>Eventually I settled on tessellation as a decent proxy for physics. The technique is quite simple in concept.</p>

<p>First, curves are flattened into line segments with configurable accuracy (aka <em>“tolerance”</em>).</p>

<p><img src="/blog/assets/Flattening.svg" alt="Tessellation" /><br />
<em>Attribution: Image from <a href="https://github.com/nical/lyon">Lyon</a>, licensed under MIT/Apache 2.0</em></p>

<p>Then, vertices are paired to generate a triangle mesh. The resulting triangle mesh is used to derive a <a href="https://en.wikipedia.org/wiki/Convex_hull_algorithms">convex hull</a> hitbox.</p>

<p><img src="/blog/assets/Tessellation.svg" alt="Tessellation" /><br />
<em>Attribution: Image from <a href="https://github.com/nical/lyon">Lyon</a>, licensed under MIT/Apache 2.0. Modified by Spencer Imbleau.</em></p>

<p>This is obviously not true collision detection between arbitrary curves with infinitessimal precision, but was rather a compromise given a lack of algorithmically efficient collision detection between curves.</p>

<p>As with most things in game development, this is a hack, but a good one! Even when using a liberal tolerance for curve flattening, physics yield a high level of collision fidelity, nearly indistinguishable from the true underlying curves. Granted, this only works great since my egg is convex in nature.</p>

<p><img src="/blog/assets/vong-closeup.png" alt="Vong Closeup" /></p>

<p>The <a href="https://github.com/nical/lyon"><code class="language-plaintext highlighter-rouge">lyon</code></a> library was used for tessellation and <a href="https://rapier.rs/"><code class="language-plaintext highlighter-rouge">bevy_rapier</code></a> was used for collision detection, linear velocity, and angular momentum.</p>

<h3 id="demo">Demo</h3>

<p><strong>Is the frame immediately below solid black?</strong> Vong uses compute shaders. Make sure your browser is updated to <a href="https://chromestatus.com/feature/6213121689518080">Chrome M113</a> or another browser compatible with <a href="https://caniuse.com/?search=webgpu">WebGPU</a>!<br />
<em>Edit 1/31: It seems Firefox Nightly, even with the WebGPU flag, is having issues. For now, please use Chrome.</em></p>

<iframe width="100%" height="400px" src="https://nuzzles.github.io/vong/" frameborder="1px solid black" style="background-color: black" allowfullscreen=""></iframe>

<p>Controls:</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">Up</code>/<code class="language-plaintext highlighter-rouge">Down</code> (Arrow keys): Scale camera</li>
  <li><code class="language-plaintext highlighter-rouge">PgUp</code>/<code class="language-plaintext highlighter-rouge">PgDown</code>/<code class="language-plaintext highlighter-rouge">Home</code>/<code class="language-plaintext highlighter-rouge">End</code>: Free camera</li>
  <li><code class="language-plaintext highlighter-rouge">W</code>/<code class="language-plaintext highlighter-rouge">S</code>: Left bacon</li>
  <li><code class="language-plaintext highlighter-rouge">I</code>/<code class="language-plaintext highlighter-rouge">K</code>: Right bacon</li>
  <li><code class="language-plaintext highlighter-rouge">C</code>: Watch egg intensely</li>
</ul>

<p>You may find the source code <a href="https://github.com/nuzzles/vong">here</a>. You may also build and play this natively with <code class="language-plaintext highlighter-rouge">cargo run</code>.</p>

<h2 id="what-now">What now?</h2>

<p>I look forward to following up with future development on vector games. I’m committing this year to publishing more about where this is all leading, but for now, please keep in touch. The next blog planned is on WebRTC to bring our games to life with multiplayer.</p>]]></content><author><name>Spencer Imbleau</name></author><category term="rust" /><category term="graphics" /><summary type="html"><![CDATA[Lessons from making the first modern compute-centric vector graphic game for web]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://spencer.imbleau.com/og-banner.png" /><media:content medium="image" url="https://spencer.imbleau.com/og-banner.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">NVIDIA GPU profiling with Rust</title><link href="https://spencer.imbleau.com/blog/rust/graphics/2022/06/19/nvidia-gpu-profiling-with-rust.html" rel="alternate" type="text/html" title="NVIDIA GPU profiling with Rust" /><published>2022-06-19T00:00:00+00:00</published><updated>2022-06-19T00:00:00+00:00</updated><id>https://spencer.imbleau.com/blog/rust/graphics/2022/06/19/nvidia-gpu-profiling-with-rust</id><content type="html" xml:base="https://spencer.imbleau.com/blog/rust/graphics/2022/06/19/nvidia-gpu-profiling-with-rust.html"><![CDATA[<p>Not too long ago, I had to perform CPU and GPU profiling in support of rendering a thesis<sup id="fnref:thesis"><a href="#fn:thesis" class="footnote" rel="footnote" role="doc-noteref">1</a></sup>. When I realized the ecosystem of GPU profiling tools in Rust was somewhat immature, I had to scaffold my own tools out of necessity. This post will demonstrate how anyone can use the NVIDIA Tools Extension SDK (NVTX) from Rust to make GPU and CPU profiling trivial.</p>

<h2 id="introduction-to-nvtx">Introduction to NVTX</h2>

<p>NVTX is a C-based API for annotating events, code, and resources in your application. These annotations get captured by NVIDIA profiling software, such as NSight Systems<sup id="fnref:nsight_systems"><a href="#fn:nsight_systems" class="footnote" rel="footnote" role="doc-noteref">2</a></sup>. Emitting tracer annotations during the runtime of your program can help identify key issues such as hardware starvation, insufficient parallelization, expensive algorithms, and more. Later in this post we will explore an example (<a href="#show-me">#show-me</a>).</p>

<p>Thankfully, one of Rust’s promises is its ability to interoperate with C APIs through foreign function interfacing (FFI) at identical performance<sup id="fnref:zero_cost_abstractions"><a href="#fn:zero_cost_abstractions" class="footnote" rel="footnote" role="doc-noteref">3</a></sup>. This is called a zero-cost abstraction, and allows us to call NVTX functions with identical performance to C, in Rust.</p>

<h2 id="nsight-systems">NSight Systems</h2>

<p>NSight Systems is a feature-ruch CLI and GUI profiler which executes a command and samples certain opt-in measurements you subscribe to. For example, if I instruct NVIDIA Nsight Systems to run a binary, Nsight Systems will launch the binary on my behalf, collect measurements while waiting for termination, and deliver a report file. In that order.</p>

<p><img src="/blog/assets/NSight_Systems_Example.png" alt="NSight Systems" /></p>

<p>These reports can be better understood with annotations from my <a href="https://crates.io/crates/nvtx">nvtx crate</a>. With it, you may name your threads, add markers, and annotate timespans. Tracers emitted during program runtime appear on the NVTX layer of a report file when capturing NVTX annotations. The following image shows an example of some tracers, with nothing else captured.</p>

<p><img src="/blog/assets/NVTX_Example.png" alt="NSight Annotations" /></p>

<h2 id="code">Code</h2>

<p>The following sections provide some examples which demonstrate how trivial it is to augment runtime with annotations.</p>

<h3 id="markers">Markers</h3>

<p>Firstly, I will show markers. Markers are used to tag instantaneous events during the execution of an application. For example, dropping a marker can be used to annotate a step in an algorithm.</p>

<p>Markers are injected with the <code class="language-plaintext highlighter-rouge">mark!</code> macro, which accepts a name for the event. It behaves similarly to the <code class="language-plaintext highlighter-rouge">println!</code> macro with argument formatting.</p>

<div class="language-rust highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">thread</span><span class="p">::</span><span class="n">sleep</span><span class="p">;</span>
<span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">time</span><span class="p">::</span><span class="n">Duration</span><span class="p">;</span>

<span class="k">use</span> <span class="nn">nvtx</span><span class="p">::</span><span class="n">mark</span><span class="p">;</span>

<span class="k">fn</span> <span class="nf">main</span><span class="p">()</span> <span class="p">{</span>
    <span class="nd">mark!</span><span class="p">(</span><span class="s">"Operation A - Begin"</span><span class="p">);</span>
    <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_millis</span><span class="p">(</span><span class="mi">150</span><span class="p">));</span> <span class="c1">// Expensive operation here</span>
    <span class="nd">mark!</span><span class="p">(</span><span class="s">"Operation B - Begin"</span><span class="p">);</span>
    <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_millis</span><span class="p">(</span><span class="mi">75</span><span class="p">));</span> <span class="c1">// Expensive operation here</span>
<span class="p">}</span>
</code></pre></div></div>

<p><img src="/blog/assets/Marker_Example.png" alt="NSight Marker" /></p>

<h3 id="thread-ranges">Thread ranges</h3>

<p>Thread ranges are similar to markers, except they annotate a span of time, rather than an instant of time. For example, a thread range can track the total time of an algorithm or process.</p>

<p>Ranges are pushed with the <code class="language-plaintext highlighter-rouge">range_push!</code> macro, and popped with the <code class="language-plaintext highlighter-rouge">range_pop!</code> macro. Thread ranges are also safe to push and pop over thread boundaries, hence <em>thread</em> in the name.</p>

<div class="language-rust highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">thread</span><span class="p">::</span><span class="n">sleep</span><span class="p">;</span>
<span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">time</span><span class="p">::</span><span class="n">Duration</span><span class="p">;</span>

<span class="k">use</span> <span class="nn">nvtx</span><span class="p">::{</span><span class="n">range_pop</span><span class="p">,</span> <span class="n">range_push</span><span class="p">};</span>

<span class="k">fn</span> <span class="nf">main</span><span class="p">()</span> <span class="p">{</span>
    <span class="nd">range_push!</span><span class="p">(</span><span class="s">"My Range"</span><span class="p">);</span>
    <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_millis</span><span class="p">(</span><span class="mi">100</span><span class="p">));</span> <span class="c1">// Expensive operation here</span>
    <span class="nd">range_pop!</span><span class="p">();</span>
<span class="p">}</span>
</code></pre></div></div>

<p><img src="/blog/assets/Range_Example.png" alt="NVTX Thread Range" /></p>

<h3 id="thread-naming">Thread naming</h3>

<p>Lastly, threads can be named with the <code class="language-plaintext highlighter-rouge">name_thread!</code> macro. This macros will alias the calling OS thread with a friendly name. Currently, this macro works cross-platform and has been tested on MacOS, Windows, Linux, and Android.</p>

<div class="language-rust highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">thread</span><span class="p">::</span><span class="n">sleep</span><span class="p">;</span>
<span class="k">use</span> <span class="nn">std</span><span class="p">::</span><span class="nn">time</span><span class="p">::</span><span class="n">Duration</span><span class="p">;</span>

<span class="k">use</span> <span class="nn">nvtx</span><span class="p">::</span><span class="n">name_thread</span><span class="p">;</span>

<span class="k">fn</span> <span class="nf">main</span><span class="p">()</span> <span class="p">{</span>
    <span class="nd">name_thread!</span><span class="p">(</span><span class="s">"Main Thread"</span><span class="p">);</span>
    <span class="k">let</span> <span class="n">handler2</span> <span class="o">=</span> <span class="nn">thread</span><span class="p">::</span><span class="nf">spawn</span><span class="p">(||</span> <span class="p">{</span>
        <span class="nd">name_thread!</span><span class="p">(</span><span class="s">"Thread 2"</span><span class="p">);</span>
        <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_secs_f32</span><span class="p">(</span><span class="mf">0.2</span><span class="p">));</span>
    <span class="p">});</span>
    <span class="k">let</span> <span class="n">handler3</span> <span class="o">=</span> <span class="nn">thread</span><span class="p">::</span><span class="nf">spawn</span><span class="p">(||</span> <span class="p">{</span>
        <span class="nd">name_thread!</span><span class="p">(</span><span class="s">"Thread 3"</span><span class="p">);</span>
        <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_secs_f32</span><span class="p">(</span><span class="mf">0.3</span><span class="p">));</span>
    <span class="p">});</span>
    <span class="nf">sleep</span><span class="p">(</span><span class="nn">Duration</span><span class="p">::</span><span class="nf">from_secs_f32</span><span class="p">(</span><span class="mf">0.5</span><span class="p">));</span>
    <span class="n">handler2</span><span class="nf">.join</span><span class="p">()</span><span class="nf">.unwrap</span><span class="p">();</span>
    <span class="n">handler3</span><span class="nf">.join</span><span class="p">()</span><span class="nf">.unwrap</span><span class="p">();</span>
<span class="p">}</span>
</code></pre></div></div>

<p><img src="/blog/assets/Thread_Example.png" alt="NVTX Thread Naming" /></p>

<h2 id="show-me">Show me</h2>

<p>To demonstrate NSight Systems, NVTX, and Rust, I will present a finding I organically stumbled upon when writing a <a href="https://github.com/nuzzles/svg-tessellation-renderer">naive SVG renderer</a> using <a href="https://github.com/nical/lyon">Lyon</a> and <a href="https://wgpu.rs/">wgpu</a>.</p>

<h3 id="situation">Situation</h3>

<p>My task was assessing if tessellation could have any usefulness in optimizing vector graphics. Vector graphic renderers solve a very complicated problem; they don’t simply read rows of pixels like in raster graphics. You may imagine an SVG file as a composition of overlapping paths and lines, and a renderer must decide on varyings (scale, viewbox, etc.) to calculate the image.</p>

<p><img src="/blog/assets/renderkit.svg" alt="Sample Images" /></p>

<p>Vector tessellation makes this easier. The idea is to transmute mathematical paths, curves, and lines into discrete triangles, a friendlier format for GPUs. My renderer would cache the triangle geometry in a storage buffer and re-use the work to inexpensively redraw the SVG.</p>

<p>I chose some sample images as an organic range with varying complexity, and proceeded to benchmark the frametimes.</p>

<p><img src="/blog/assets/all_images.svg" alt="Sample Images" /></p>

<h3 id="results">Results</h3>

<p>The benchmark was simple: render 50 frames and record how long each frame took. The results are plotted below.</p>

<p><img src="/blog/assets/all_renderkit.svg" alt="My Results" /></p>

<p>Tessellation-time was not plotted here, and no frame was rendered any differently. So to my surprise, why was the first frame taking longer? Using the power of NVTX annotations, I dropped some tracers to understand the instructions temporally. I annotated the first frame as “<em>Strange Behavior</em>”.</p>

<p><img src="/blog/assets/NVTX_Trace.png" alt="NVTX Tracers" /></p>

<p>To my surprise, it seemed like the GPU wasn’t working hard until the 2nd frame. The reasoning would be that wgpu (the graphics library I used) submitted the queue of geometry in the first frame. What we are witnessing is gpu latency, specifically, the time to transfer the geometry to the GPU storage buffers and return. This became an obvious que when able to inspect and see a timeline with ques: blocked state, PCIe bandwidth, and GPU occupancy.</p>

<p>This is one example of how annotations have helped me, and work will continue to help in more ways.</p>

<h2 id="features-in-progress">Features in progress</h2>

<p>I have only ported a fraction of the NVTX SDK, and there are some noteworthy features outstanding. NVTX can be used to measure CPU code blocks, track lifetime of CPU resources (e.g., malloc), log critical events<sup id="fnref:nvtx_def"><a href="#fn:nvtx_def" class="footnote" rel="footnote" role="doc-noteref">4</a></sup>, and more. I’d love to port <a href="https://nvidia.github.io/NVTX/doxygen/index.html#DOMAINS">filtering tracers by domains</a> and <a href="https://nvidia.github.io/NVTX/doxygen/index.html#RESOURCE_NAMING">resource naming</a>. Feel free to contribute or check out the <a href="https://github.com/nuzzles/nvtx/issues">issue board</a>.</p>

<hr />

<!-- Note: There must be a blank line between every two lines of the footnote difinition.  -->

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:thesis">
      <p><a href="http://dx.doi.org/10.13140/RG.2.2.25593.54887">Understanding Hardware-Accelerated 2D Vector Graphics</a> <a href="#fnref:thesis" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:nsight_systems">
      <p><a href="https://developer.nvidia.com/nsight-systems">NVIDIA NSight Systems</a> <a href="#fnref:nsight_systems" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:zero_cost_abstractions">
      <p><a href="https://blog.rust-lang.org/2015/04/24/Rust-Once-Run-Everywhere.html">Run Once, Run Everywhere</a> <a href="#fnref:zero_cost_abstractions" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:nvtx_def">
      <p><a href="https://docs.nvidia.com/nsight-visual-studio-edition/2020.1/nvtx/index.html#nvidia-tools-extension-library-nvtx">The NVIDIA Tools Extension Library (NVTX)</a> <a href="#fnref:nvtx_def" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Spencer Imbleau</name></author><category term="rust" /><category term="graphics" /><summary type="html"><![CDATA[The NVIDIA Tools Exension SDK (NVTX) has a lot to offer for GPU and CPU profiling. This blog shows how to leverage these tools with Rust, and how we can make sense of the results.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://spencer.imbleau.com/og-banner.png" /><media:content medium="image" url="https://spencer.imbleau.com/og-banner.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Hello world</title><link href="https://spencer.imbleau.com/blog/2021/05/30/hello-world.html" rel="alternate" type="text/html" title="Hello world" /><published>2021-05-30T00:00:00+00:00</published><updated>2021-05-30T00:00:00+00:00</updated><id>https://spencer.imbleau.com/blog/2021/05/30/hello-world</id><content type="html" xml:base="https://spencer.imbleau.com/blog/2021/05/30/hello-world.html"><![CDATA[<p>Hello, I’m going to be using this space to collect my thoughts and share them.</p>

<p>I will especially do so in my favorite topics, such as open source, technical subjects, and lifestyle (as a programmer).</p>]]></content><author><name>Spencer Imbleau</name></author><summary type="html"><![CDATA[The "O brave new world" of someone less poetic.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://spencer.imbleau.com/og-banner.png" /><media:content medium="image" url="https://spencer.imbleau.com/og-banner.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>