<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="/feed.xml" rel="self" type="application/atom+xml" /><link href="/" rel="alternate" type="text/html" /><updated>2026-01-31T07:44:14+00:00</updated><id>/feed.xml</id><title type="html">Linzhi Wu’s Blog</title><entry><title type="html">A Practical RTC-Style Stitching Workaround</title><link href="/jekyll/update/2026/01/26/implementing-rtc-correctly.html" rel="alternate" type="text/html" title="A Practical RTC-Style Stitching Workaround" /><published>2026-01-26T03:41:56+00:00</published><updated>2026-01-26T03:41:56+00:00</updated><id>/jekyll/update/2026/01/26/implementing-rtc-correctly</id><content type="html" xml:base="/jekyll/update/2026/01/26/implementing-rtc-correctly.html"><![CDATA[<h2 id="introduction">Introduction</h2>

<p>Recent vision-language-action (VLA) models have shown impressive capability in robotic manipulation, but their high inference latency remains a major obstacle for real-time deployment. The Real-Time Chunking (RTC) framework addresses this issue by enabling asynchronous execution of action-chunked policies, reformulating real-time control as an inference-time inpainting problem. Without any retraining, RTC allows a policy to generate future action chunks while safely executing previously committed actions, achieving smooth and delay-robust behavior in both simulation and real-world robotic tasks.</p>

<p>While reproducing the results of this work and experimenting with its released implementation, I ran into a gradient-handling pitfall that can meaningfully affect cross-chunk continuity and runtime behavior. Specifically, when implementing RTC-style per-step “stitching” guidance, it is tempting to let gradients propagate through the learned vector field / flow policy \(v(\cdot)\) via automatic differentiation.</p>

<p>In retrospect, the approach described in this post is best viewed as a <strong>pragmatic workaround</strong>: it often produces an RTC-like continuity effect, but it is not necessarily a principled or faithful implementation of RTC in a strict algorithmic sense. The goal here is to document an implementation choice—stopping gradients through \(v\)—that can strengthen the guidance signal in practice, and to show what changes when you do (or do not) take that shortcut.</p>

<h2 id="preliminaries">Preliminaries</h2>

<h3 id="action-chunking-and-horizons">Action chunking and horizons</h3>

<p>RTC considers an action-chunking policy \(\pi(A_t \mid o_t)\), where
\(A_t = [a_t, a_{t+1}, \ldots, a_{t+H-1}]\) is a chunk of \(H\) future actions (the <strong>prediction horizon</strong>).<br />
At rollout time, only the first \(s \le H\) actions are executed (the <strong>execution horizon</strong>). Longer \(s\) improves temporal consistency but reduces reactivity; shorter \(s\) increases the chance of discontinuities between chunks.</p>

<h3 id="flow-policy-sampling">Flow policy sampling</h3>

<p>RTC focuses on policies trained with conditional flow matching (diffusion policies can be converted to flow policies at inference time).<br />
A chunk is generated by sampling Gaussian noise and integrating the learned velocity field \(v_\pi\) over flow time \(\tau \in [0,1)\) with \(n\) denoising steps:</p>

\[A^{\tau + \frac{1}{n}}_t = A^\tau_t + \frac{1}{n}\, v_\pi(A^\tau_t, o_t, \tau).
\tag{1}\]

<h3 id="real-time-constraint-and-inference-delay">Real-time constraint and inference delay</h3>

<p>Let \(\Delta t\) be the controller period and \(\delta\) the wall-clock time to generate one chunk. RTC defines the <strong>inference delay</strong></p>

\[d := \left\lfloor \frac{\delta}{\Delta t} \right\rfloor,
\tag{2}\]

<p>the number of control steps between receiving \(o_t\) and having \(A_t\) available. 
A naive asynchronous strategy can be real-time when \(d \le H - s\), but it may cause a discontinuous switch at the chunk boundary because the policy cannot predict what happens during the delay window.</p>

<h3 id="rtc-as-inpainting-with-soft-masking">RTC as inpainting with soft masking</h3>

<p>RTC frames asynchronous chunking as an <strong>inpainting</strong> problem: while executing the previous chunk, the actions that are guaranteed to be executed due to delay are treated as “frozen,” and the rest of the new chunk is generated to be compatible with the previous one.<br />
To strengthen cross-chunk continuity, RTC uses <strong>soft masking</strong> weights \(W\): the first \(d\) overlapping actions receive weight 1, the non-overlapping tail receives weight 0, and the intermediate region decays smoothly from 1 to 0. The guided inference step is implemented via a vector–Jacobian product (VJP) computed with reverse-mode autodifferentiation, which is exactly where gradient-handling becomes subtle in practice.</p>

<h2 id="the-πgdm-correction-term-and-a-stop-gradient-approximation">The ΠGDM Correction Term and a Stop-Gradient Approximation</h2>

<p>In RTC, the ΠGDM-style corrected vector field can be written as</p>

\[v_{\Pi\mathrm{GDM}}(A_t^\tau,o_t,\tau)
=
v(A_t^\tau,o_t,\tau)
+
\min\!\left(
\beta,\;
\frac{1-\tau}{\tau \cdot r_\tau^{2}}
\right)
\bigl(Y - \hat{A}_t^{1}\bigr)^\top \mathrm{diag}(W)\;
\frac{\partial \hat{A}_t^{1}}{\partial A_t^\tau},
\tag{3}\]

<p>where</p>

\[\hat{A}_t^{1}=A_t^\tau+(1-\tau)\,v(A_t^\tau,o_t,\tau).
\tag{4}\]

<p>If we apply the chain rule strictly, then</p>

\[\frac{\partial \hat{A}_t^{1}}{\partial A_t^\tau}
=
I + (1-\tau)\frac{\partial v(A_t^\tau,o_t,\tau)}{\partial A_t^\tau}.
\tag{5}\]

<p>At first glance, it seems natural to use (5) when computing the VJP in (3).</p>

<p>However, if your goal is a “residual-style” stitching constraint (the inpainting/continuity view), a common practical approximation is to <strong>treat \(v\) as constant</strong> for the VJP (i.e., stop gradients through \(v\)). Under this approximation,</p>

\[\frac{\partial \hat{A}_t^{1}}{\partial A_t^\tau} = I,
\tag{6}\]

<p>and the correction direction reduces to the most direct weighted residual:</p>

\[\min\!\left(
\beta,\;
\frac{1-\tau}{\tau \cdot r_\tau^{2}}
\right)
\bigl(Y - \hat{A}_t^{1}\bigr)^\top \mathrm{diag}(W).
\tag{7}\]

<h3 id="why-stop-gradients-through-v-in-practice">Why stop gradients through \(v\) in practice?</h3>

<p>The key point is that this guidance step is trying to enforce <strong>cross-chunk continuity</strong>—conceptually an inpainting constraint that pulls the current prediction toward a target \(Y\) in the overlap region (weighted by \(W\)). In practice, letting gradients flow “through the network” can sometimes <strong>dampen the effective guidance signal</strong>, making the resulting correction term too small to meaningfully steer the trajectory at chunk boundaries. When the guidance is underpowered, the sampler largely follows the unconditional flow, and cross-chunk continuity can degrade because the constraint is not enforced strongly enough.</p>

<p>Here is a minimal JAX sketch of the stop-gradient workaround for 1→0 flow integration (e.g., \(\pi_{0}\)). In <code class="language-plaintext highlighter-rouge">denoiser()</code>, we apply <code class="language-plaintext highlighter-rouge">jax.lax.stop_gradient(v_t)</code> so that <code class="language-plaintext highlighter-rouge">jax.vjp</code> does not backpropagate through <code class="language-plaintext highlighter-rouge">v_t_fn</code> (i.e., treats <code class="language-plaintext highlighter-rouge">v_t</code> as constant w.r.t. <code class="language-plaintext highlighter-rouge">x_t</code>).</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kn">import</span> <span class="nn">jax</span>
<span class="kn">import</span> <span class="nn">jax.numpy</span> <span class="k">as</span> <span class="n">jnp</span>
<span class="kn">from</span> <span class="nn">typing</span> <span class="kn">import</span> <span class="n">Callable</span>


<span class="k">def</span> <span class="nf">pinv_corrected_velocity</span><span class="p">(</span>
    <span class="n">v_t_fn</span><span class="p">:</span> <span class="n">Callable</span><span class="p">,</span> <span class="c1"># ([ah ad], float) -&gt; [ah ad]
</span>    <span class="n">x_t</span><span class="p">:</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">,</span> <span class="c1"># [b ah ad]
</span>    <span class="n">t</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span>
    <span class="n">prefix_actions</span><span class="p">:</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">,</span> <span class="c1"># [b ah ad]
</span>    <span class="n">inference_delay</span><span class="p">:</span> <span class="nb">int</span><span class="p">,</span>
    <span class="n">prefix_attention_horizon</span><span class="p">:</span> <span class="nb">int</span><span class="p">,</span>
    <span class="n">max_guidance_weight</span><span class="p">:</span> <span class="nb">float</span><span class="p">,</span>
<span class="p">)</span> <span class="o">-&gt;</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">:</span> <span class="c1"># [b ah ad]
</span>    <span class="o">@</span><span class="n">jax</span><span class="p">.</span><span class="n">vmap</span>
    <span class="k">def</span> <span class="nf">_pinv_corrected_velocity</span><span class="p">(</span>
        <span class="n">x_t</span><span class="p">:</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">,</span> <span class="c1"># [ah ad]
</span>        <span class="n">y</span><span class="p">:</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">,</span> <span class="c1"># [ah ad]
</span>    <span class="p">)</span> <span class="o">-&gt;</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">:</span> <span class="c1"># [ah ad]
</span>        <span class="k">def</span> <span class="nf">denoiser</span><span class="p">(</span><span class="n">x_t</span><span class="p">:</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="nb">tuple</span><span class="p">[</span><span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">,</span> <span class="n">jax</span><span class="p">.</span><span class="n">Array</span><span class="p">]:</span>
            <span class="n">v_t</span> <span class="o">=</span> <span class="n">v_t_fn</span><span class="p">(</span><span class="n">x_t</span><span class="p">,</span> <span class="n">t</span><span class="p">)</span>
            <span class="n">x_0</span> <span class="o">=</span> <span class="n">x_t</span> <span class="o">-</span> <span class="n">jax</span><span class="p">.</span><span class="n">lax</span><span class="p">.</span><span class="n">stop_gradient</span><span class="p">(</span><span class="n">v_t</span><span class="p">)</span> <span class="o">*</span> <span class="n">t</span>
            <span class="k">return</span> <span class="n">x_0</span><span class="p">,</span> <span class="n">v_t</span>

        <span class="n">x_0</span><span class="p">,</span> <span class="n">vjp_fun</span><span class="p">,</span> <span class="n">v_t</span> <span class="o">=</span> <span class="n">jax</span><span class="p">.</span><span class="n">vjp</span><span class="p">(</span><span class="n">denoiser</span><span class="p">,</span> <span class="n">x_t</span><span class="p">,</span> <span class="n">has_aux</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
        <span class="n">error</span> <span class="o">=</span> <span class="p">(</span><span class="n">y</span> <span class="o">-</span> <span class="n">x_0</span><span class="p">)</span> <span class="o">*</span> <span class="n">get_prefix_weights</span><span class="p">(</span><span class="n">inference_delay</span><span class="p">,</span> <span class="n">prefix_attention_horizon</span><span class="p">,</span> <span class="n">prefix_actions</span><span class="p">.</span><span class="n">shape</span><span class="p">[</span><span class="mi">1</span><span class="p">])[:,</span> <span class="bp">None</span><span class="p">]</span>
        <span class="n">pinv_correction</span> <span class="o">=</span> <span class="n">vjp_fun</span><span class="p">(</span><span class="n">error</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span>
        <span class="c1"># constants from paper
</span>        <span class="n">inv_r2</span> <span class="o">=</span> <span class="p">(</span><span class="n">t</span><span class="o">**</span><span class="mi">2</span> <span class="o">+</span> <span class="p">(</span><span class="mi">1</span> <span class="o">-</span> <span class="n">t</span><span class="p">)</span> <span class="o">**</span> <span class="mi">2</span><span class="p">)</span> <span class="o">/</span> <span class="p">(</span><span class="n">t</span><span class="o">**</span><span class="mi">2</span><span class="p">)</span>
        <span class="n">c</span> <span class="o">=</span> <span class="n">jnp</span><span class="p">.</span><span class="n">nan_to_num</span><span class="p">(</span><span class="n">t</span> <span class="o">/</span> <span class="p">(</span><span class="mi">1</span> <span class="o">-</span> <span class="n">t</span><span class="p">),</span> <span class="n">posinf</span><span class="o">=</span><span class="n">max_guidance_weight</span><span class="p">)</span>
        <span class="n">guidance_weight</span> <span class="o">=</span> <span class="n">jnp</span><span class="p">.</span><span class="n">minimum</span><span class="p">(</span><span class="n">c</span> <span class="o">*</span> <span class="n">inv_r2</span><span class="p">,</span> <span class="n">max_guidance_weight</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">v_t</span> <span class="o">-</span> <span class="n">guidance_weight</span> <span class="o">*</span> <span class="n">pinv_correction</span>

    <span class="k">return</span> <span class="n">_pinv_corrected_velocity</span><span class="p">(</span><span class="n">x_t</span><span class="p">,</span> <span class="n">prefix_actions</span><span class="p">)</span>
</code></pre></div></div>

<h2 id="experiment-what-does-this-rtc-like-workaround-look-like-in-practice">Experiment: What does this RTC-like workaround look like in practice?</h2>

<p>We fine-tuned \(\pi_{05}\) on a dual-arm Flexiv robot to perform a LEGO sorting task. The figure below shows results from this RTC-style stitching guidance (using the stop-gradient workaround), evaluated on dataset trajectories.</p>

<ul>
  <li><span style="color: purple;">purple</span> denotes the actions generated with the RTC-style stitching guidance (stop-gradient workaround),</li>
  <li><span style="color: green;">green</span> denotes the corresponding actions from the previous action chunk used for guidance,</li>
  <li><span style="color: red;">red dashed</span> denotes the actions generated without this stitching guidance,</li>
  <li><span style="color: gray;">gray dashed</span> denotes the inference delay.</li>
</ul>

<p>We use an action chunk length of 100. The RTC-related parameters in this test are: <code class="language-plaintext highlighter-rouge">inference_delay = 16</code>, <code class="language-plaintext highlighter-rouge">max_guidance_weight = 10</code>, and <code class="language-plaintext highlighter-rouge">execution_horizon = 64</code>.</p>

<p>Notably, the y-axis is measured in meters. The 14 dimensions in the plot correspond to \((x, y, z, r_x, r_y, r_z, \text{gripper}) \times 2\)</p>

<p><img src="/img/correction,max_guidance=10.png" alt="" /></p>

<p>In contrast, using the same parameters and input data but simply removing <code class="language-plaintext highlighter-rouge">jax.lax.stop_gradient(v_t)</code> yields the following result:</p>

<p><img src="/img/original,max_guidance=10.png" alt="" /></p>

<p>In addition, tuning <code class="language-plaintext highlighter-rouge">max_guidance_weight</code> suggests that <code class="language-plaintext highlighter-rouge">max_guidance_weight = 10</code> is better suited for the default \(\pi_{05}\) implementation. This smoothing idea is also discussed in Soare’s blog post.<sup id="fnref:soare2025" role="doc-noteref"><a href="#fn:soare2025" class="footnote" rel="footnote">1</a></sup></p>

<h3 id="max_guidance_weight--5">max_guidance_weight = 5</h3>

<p><img src="/img/correction,max_guidance_weight=5.png" alt="" /></p>

<h3 id="max_guidance_weight--100">max_guidance_weight = 100</h3>

<p><img src="/img/correction,max_guidance_weight=100.png" alt="" /></p>

<h3 id="unmodified-rtc-max_guidance--5">Unmodified RTC max_guidance = 5</h3>
<p><img src="/img/original,max_guidance=5.png" alt="" /></p>

<h3 id="unmodified-rtc-max_guidance--100">Unmodified RTC max_guidance = 100</h3>
<p><img src="/img/original,max_guidance=100.png" alt="" /></p>

<blockquote>
  <p>⚠️ <strong>Important note / Limitations.</strong><br />
These experiments were conducted under fairly limited conditions—only a small number of tasks and just a handful of episodes—though we tested multiple action parameterizations, including absolute end-effector poses, relative end-effector poses, relative joint angles, and absolute joint angles. This method may encounter issues in certain scenarios, and it is only provided as a reference.</p>
</blockquote>

<h2 id="citation">Citation</h2>

<p>Please feel free to cite this work as</p>

<div class="language-bibtex highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nc">@misc</span><span class="p">{</span> <span class="nl">wu2026rtc</span><span class="p">,</span>
  <span class="na">author</span> <span class="p">=</span> <span class="s">{ Linzhi Wu and Junjie Xu }</span><span class="p">,</span>
  <span class="na">title</span>  <span class="p">=</span> <span class="s">{ A Practical RTC-Style Stitching Workaround }</span><span class="p">,</span>
  <span class="na">year</span>   <span class="p">=</span> <span class="s">{ 2026 }</span><span class="p">,</span>
  <span class="na">url</span>    <span class="p">=</span> <span class="s">{ https://xiaozhisky1.github.io/jekyll/update/2026/01/26/implementing-rtc-correctly.html }</span><span class="p">,</span>
  <span class="na">note</span>   <span class="p">=</span> <span class="s">{ Blog post }</span>
<span class="p">}</span>
</code></pre></div></div>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:soare2025" role="doc-endnote">
      <p>Alexander Soare, <a href="https://alexander-soare.github.io/robotics/2025/08/05/smooth-as-butter-robot-policies.html"><em>Smooth-As-Butter Robot Policies</em></a>, 2025 (blog post). <a href="#fnref:soare2025" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Linzhi Wu, Junjie Xu</name></author><category term="jekyll" /><category term="update" /><summary type="html"><![CDATA[Introduction]]></summary></entry></feed>