<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://arte-lab.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://arte-lab.github.io/" rel="alternate" type="text/html" /><updated>2026-07-10T02:07:52+00:00</updated><id>https://arte-lab.github.io/feed.xml</id><title type="html">ARTE Group</title><subtitle>ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.</subtitle><entry><title type="html">Mingrui Wu presented OpenBench at CVPR 2026</title><link href="https://arte-lab.github.io/news/2026/06/07/mingrui-wu-cvpr-2026.html" rel="alternate" type="text/html" title="Mingrui Wu presented OpenBench at CVPR 2026" /><published>2026-06-07T00:00:00+00:00</published><updated>2026-06-07T00:00:00+00:00</updated><id>https://arte-lab.github.io/news/2026/06/07/mingrui-wu-cvpr-2026</id><content type="html" xml:base="https://arte-lab.github.io/news/2026/06/07/mingrui-wu-cvpr-2026.html"><![CDATA[<p>From June 3 to June 7, 2026, <a href="https://arte-lab.github.io/members/Mingrui-Wu.html">Mingrui Wu</a> attended the <a href="https://cvpr.thecvf.com/Conferences/2026">IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026)</a> at the Colorado Convention Center in Denver, Colorado. During the conference, he presented our paper, <a href="/research/example-project/"><em>From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs</em></a> (<a href="https://arxiv.org/abs/2512.19683">arXiv</a>).</p>

<p>In this work, we introduce OpenBench, a benchmark for evaluating the spatial reasoning ability of multimodal large language models in realistic open-world scenes. Built from pedestrian-perspective multimodal data with metric 3D cues, OpenBench covers question types ranging from qualitative relations to quantitative and kinematic reasoning. Our results show that performance gains reported in indoor settings do not transfer directly to open-world scenarios, highlighting an important gap in grounded spatial intelligence.</p>

<p>Congratulations to Mingrui and the team on sharing this work with the CVPR 2026 community!</p>

<div class="gallery" data-fit="false" data-flat="false" data-number="1"><a href="/research/example-project/" class="gallery_item" data-tooltip="Mingrui Wu at CVPR 2026">
    <img src="/news-imgs/cvpr2026-mingrui/2026-6-3_0.jpg" />
  </a></div>]]></content><author><name>Mingrui Wu</name></author><category term="news" /><category term="conference" /><category term="CVPR" /><category term="MLLM" /><category term="spatial reasoning" /><category term="OpenBench" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">Yanhao Wu introduces AlignDrive for coordinated autonomous-driving planning</title><link href="https://arte-lab.github.io/news/2026/05/27/yanhao-wu-align-drive.html" rel="alternate" type="text/html" title="Yanhao Wu introduces AlignDrive for coordinated autonomous-driving planning" /><published>2026-05-27T00:00:00+00:00</published><updated>2026-05-27T00:00:00+00:00</updated><id>https://arte-lab.github.io/news/2026/05/27/yanhao-wu-align-drive</id><content type="html" xml:base="https://arte-lab.github.io/news/2026/05/27/yanhao-wu-align-drive.html"><![CDATA[<p>Congratulations to the Team on Achieving State-of-the-Art in Autonomous Driving! We are thrilled to announce that our recent joint work with Horizon Robotics, led by <a href="https://arte-lab.github.io/members/Yanhao-Wu.html">Yanhao Wu</a>, has achieved Top-1 performance on both the Bench2drive and the newly released <a href="https://simonger.github.io/fail2drive/">Fail2Drive official leaderboard</a>. Notably, our approach demonstrated exceptional resilience, proving to have the best generalization ability among competing methods.</p>

<p>Discover how we did it in our paper: <a href="/research/aligndrive/"><em>AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving</em></a> (<a href="https://arxiv.org/abs/2601.01762">arXiv</a>). For a deeper dive into the methodology and to view our qualitative results, please visit our official <a href="https://yanhaowu.github.io/AlignDrive/">project homepage</a>.</p>

<p>AlignDrive takes aim at a subtle but important planning mismatch in autonomous driving: where to go and how fast to move should be decided together. The method conditions longitudinal prediction on the planned drive path, so speed reasoning becomes tied to the vehicle’s lateral choices and surrounding-agent interactions. It also uses planning-focused augmentation to expose the model to rare safety-critical cases, improving robustness on challenging closed-loop benchmarks.</p>

<div class="gallery" data-fit="true" data-flat="false" data-number="2"><a class="gallery_item" data-tooltip="AlignDrive teaser">
    <img src="/news-imgs/aligndrive-2026/teaser.png" />
  </a><a class="gallery_item" data-tooltip="AlignDrive method overview">
    <img src="/news-imgs/aligndrive-2026/method.png" />
  </a></div>

<div class="gallery" data-fit="true" data-flat="false" data-number="1"><a class="gallery_item" data-tooltip="AlignDrive results">
    <img src="/news-imgs/aligndrive-2026/results.jpg" />
  </a></div>]]></content><author><name>Yanhao Wu</name></author><category term="news" /><category term="autonomous driving" /><category term="end-to-end planning" /><category term="arXiv" /><category term="AlignDrive" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">Tong Zhang and Zhaozhi Wang joined CICV 2026</title><link href="https://arte-lab.github.io/news/2026/05/22/tong-zhang-zhaozhi-wang-cicv-2026.html" rel="alternate" type="text/html" title="Tong Zhang and Zhaozhi Wang joined CICV 2026" /><published>2026-05-22T00:00:00+00:00</published><updated>2026-05-22T00:00:00+00:00</updated><id>https://arte-lab.github.io/news/2026/05/22/tong-zhang-zhaozhi-wang-cicv-2026</id><content type="html" xml:base="https://arte-lab.github.io/news/2026/05/22/tong-zhang-zhaozhi-wang-cicv-2026.html"><![CDATA[<p>Prof. <a href="https://arte-lab.github.io/members/Tong-Zhang.html">Tong Zhang</a>, together with team member <a href="https://arte-lab.github.io/members/Zhaozhi-Wang.html">Zhaozhi Wang</a>, was invited to participate in CICV 2026, the 13th Congress of Intelligent and Connected Vehicles Technology, held in Shanghai from May 21 to May 22, 2026.</p>

<p>During the A2 workshop, <em>World Model: A New Paradigm for Autonomous Driving Research from High-Quality Data Generation to Reinforcement Learning Training</em>, Zhaozhi delivered a talk on behalf of Prof. Zhang’s team. The presentation, titled <em>Exploring Safety for Autonomous Driving: From Scene Generation to Planning</em>, discussed how scenario generation and planning methods can support safer autonomous-driving systems. More information about the conference is available on the <a href="http://www.cicv.org.cn/CN/home/index.html">CICV official website</a>.</p>

<p>Congratulations to Prof. Zhang, our team member Zhaozhi, and the team on sharing our autonomous-driving research with the CICV community.</p>

<div class="gallery" data-fit="true" data-flat="false" data-number="3"><a class="gallery_item" data-tooltip="Zhaozhi Wang presenting at CICV 2026">
    <img src="/news-imgs/cicv/zhaozhi.jpg" />
  </a><a class="gallery_item" data-tooltip="CICV 2026 invited talk">
    <img src="/news-imgs/cicv/Zhaozhi2.jpg" />
  </a><a class="gallery_item" data-tooltip="A2 workshop group photo">
    <img src="/news-imgs/cicv/zhaozhi1.jpg" />
  </a></div>]]></content><author><name>Zhaozhi Wang</name></author><category term="news" /><category term="conference" /><category term="CICV" /><category term="autonomous driving" /><category term="world models" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">Zhaozhi Wang attended ICLR 2026 in Rio de Janeiro</title><link href="https://arte-lab.github.io/news/2026/04/27/zhaozhi-wang-attended-iclr-2026.html" rel="alternate" type="text/html" title="Zhaozhi Wang attended ICLR 2026 in Rio de Janeiro" /><published>2026-04-27T00:00:00+00:00</published><updated>2026-04-27T00:00:00+00:00</updated><id>https://arte-lab.github.io/news/2026/04/27/zhaozhi-wang-attended-iclr-2026</id><content type="html" xml:base="https://arte-lab.github.io/news/2026/04/27/zhaozhi-wang-attended-iclr-2026.html"><![CDATA[<p>From April 23 to April 27, 2026, <a href="https://arte-lab.github.io/members/Zhaozhi-Wang.html">Zhaozhi Wang</a> attended the <a href="https://iclr.cc/Conferences/2026">International Conference on Learning Representations (ICLR 2026)</a> in Rio de Janeiro, Brazil. During the conference, he presented our paper, <a href="/research/videoanchor/"><em>VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning</em></a>.</p>

<p>In this work, we identify that multimodal large language models often struggle with visual-spatial reasoning because visual tokens are overshadowed by language tokens in attention. To address this issue, we propose VideoAnchor, a plug-and-play module that reinforces shared visual cues across frames without retraining, leading to more coherent visual grounding and stronger performance on spatial reasoning benchmarks.</p>

<p>Congratulations to Zhaozhi and the group on this exciting presentation at ICLR 2026!</p>

<div class="gallery" data-fit="true" data-flat="false" data-number="3"><a class="gallery_item" data-tooltip="">
    <img src="/news-imgs/iclr2026-zhaozhi/1.jpg" />
  </a><a class="gallery_item" data-tooltip="">
    <img src="/news-imgs/iclr2026-zhaozhi/2.jpg" />
  </a><a class="gallery_item" data-tooltip="">
    <img src="/news-imgs/iclr2026-zhaozhi/3.jpg" />
  </a></div>]]></content><author><name>Zhaozhi Wang</name></author><category term="news" /><category term="conference" /><category term="ICLR" /><category term="visual-spatial reasoning" /><category term="VideoAnchor" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">The skeptic’s guide to generative AI assisted coding</title><link href="https://arte-lab.github.io/blog/2026/02/15/a-skeptics-guide-to-generative-ai-coding.html" rel="alternate" type="text/html" title="The skeptic’s guide to generative AI assisted coding" /><published>2026-02-15T00:00:00+00:00</published><updated>2026-02-15T00:00:00+00:00</updated><id>https://arte-lab.github.io/blog/2026/02/15/a-skeptics-guide-to-generative-ai-coding</id><content type="html" xml:base="https://arte-lab.github.io/blog/2026/02/15/a-skeptics-guide-to-generative-ai-coding.html"><![CDATA[<p><em>OK; it’s not really a guide, but I needed a good title</em>, this is more a story about how, and why, my perspective on AI assisted coding has evolved.</p>

<p>Since the AI-assisted coding hype began, I have been an outspoken skeptic of these technologies.  In large part, this was due to my own experience with the early models.  While I am a techie, I am very much deliberate in how I adopt and use technology.  I <em>love</em> useful tech that makes my life easier, and lets me accomplish more of my goals.  However, I <em>hate</em> bad tech; technology that wastes my time, demonstrates bad design, or diminishes rather than enriches enjoyment (yes, I know, I was quite active on Twitter for a while …. maybe that’s a post for another day).</p>

<p>When I first tried generative coding models, they simply did not produce good code.  They would sometimes produce functional code (but often not), and their solutions were almost always both messier and less efficient than what I would have done by hand.  I tasked them with implementing the projects in the classes I teach, and they either failed, or succeeded with implementations that I would hope an undergraduate at UMD would find embarrassing.  One early success, that I filed away in the back of my mind, however, was having ChatGPT help me write some highly vectorized AVX-256 code for longest common prefix computation that ended up being useful in our <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11056320/">CAPS-SA</a> work. Nonetheless, I was incredibly skeptical of people talking up (and using) these tools, and was quite outspoken about the types of mistakes they made and the limitations they had.</p>

<h2 id="a-small-series-of-successes">A small series of successes</h2>

<p>You can try to ignore the world, but you do so at your own peril.  Increasingly, colleagues and experts whose opinions I truly respect (some of whom had previously been very outspoken AI skeptics) began to express some very positive opinions.  To be clear, I think that in the tech world, there is still a <em>massive</em> overstatement of the case for and future of some of these technology. However, even some of the calmest and most judicious tech folks I follow, who had previously expressed skepticism similar to mine, began to express much more optimistic, albeit nuanced, opinions. In particular, Jon Gjengset, whose technical opinions I have come to value immensely, and whose overall judgments (and taste) I find align highly with my own, has made some very keen observations about the use of these tools, and demonstrated their utility in a way that I find that I would share.  So, I decided to give it another go.</p>

<p>Very quickly, I learned that, while I believe my opinions on the early models was correct, in that they were useful for helping those with little to no coding experience to cook up somewhat functional, but often ill-designed things, and they were often of very little use (or even a net negative) for seasoned developers. However, the current state of things is <em>entirely</em> different.  In rapid succession, I was able to obtain a series of quick and, critically, practically useful successes with the newer class of models.  I mention a few below just to give the flavor of how I was using these tools.</p>

<h3 id="mim"><code class="language-plaintext highlighter-rouge">mim</code></h3>

<p>We recently <a href="https://www.biorxiv.org/content/10.1101/2025.11.24.690271v1">pre-printed</a> a method (and associated tool) called <code class="language-plaintext highlighter-rouge">mim</code>, which is a <em>semantically aware</em> auxiliary index for gzip compressed (and block gzip compressed) FASTA and FASTQ files.  We demonstrate that with a very small auxiliary index, one can <em>massively</em> speed up the parallel decompression and parsing of these files, eliminating a bottleneck in many tasks in high throughput sequencing analysis. These indices are semantically aware (i.e. know about the structure of the read records) and so can e.g. properly synchronize between paired-end files during decompression.  <code class="language-plaintext highlighter-rouge">mim</code> is built, at a basic level, around the ideas of <code class="language-plaintext highlighter-rouge">zran</code>, and the initial implementation of <code class="language-plaintext highlighter-rouge">mim</code> was written in a mix of C and C++, with part of the indexer as a direct modification of the <code class="language-plaintext highlighter-rouge">zran</code> code. While this was useful and functional as a way to show the benefits of <code class="language-plaintext highlighter-rouge">mim</code>, I have come to greatly value keeping my software ecosystem in Rust as much as possible. It is the language  that I prefer to write, and to maintain.</p>

<p>So, I tasked Claude (Sonnet, on the free plan!) with porting the core indexing functionality of <code class="language-plaintext highlighter-rouge">mim</code> from C++ to Rust.  This was scoped to be a relatively straightforward translation, but nonetheless is something that would probably have taken me a few days to do by hand.  More critically, it would have taken me large, uninterrupted, meeting-free blocks of time.  Something of which I, as a professor, have precious little.</p>

<p>With the C++ source code (uploaded as an attachment), and a few hours worth of conversation, I was able to obtain a fully-functional Rust implementation of the <code class="language-plaintext highlighter-rouge">mim</code> indexing functionality in Rust.  It was not a “one shot” solution; there were some hiccups in how multi-part archives were handled. But each time I ran into a problem, so long as I was able to explain to Claude the technical details of the issue that I was encountering, we were able to overcome the problem.  As a result, we ended up with a <code class="language-plaintext highlighter-rouge">mim</code> implementation in Rust in short order, which has since become the core reference implementation of the idea. The current code has been polished (and parts of it rewritten by 1337 coders, like <a href="https://curiouscoding.nl/">Ragnar</a>), but getting to a functional and respectable Rust implementation, in a day, via the free tier of Sonnet, more than piqued my interest.</p>

<h3 id="in-memory-tile-collation-in-cuttlefish-1">In memory tile collation in <code class="language-plaintext highlighter-rouge">Cuttlefish</code> 1</h3>

<p>Our <a href="https://academic.oup.com/bioinformatics/article/37/Supplement_1/i177/6319696"><code class="language-plaintext highlighter-rouge">Cuttlefish</code></a> (version 1) tool still acts as the basis for building the core information that is indexed by our (<code class="language-plaintext highlighter-rouge">SSHash</code>-based) <code class="language-plaintext highlighter-rouge">piscem</code> tool — the mapper for single-cell RNA-seq and single-cell ATAC-seq data upstream of <a href="https://www.nature.com/articles/s41592-022-01408-3"><code class="language-plaintext highlighter-rouge">alevin-fry</code></a>.<br />
Cuttlefish 1 builds a compacted colored de Bruijn graph very efficiently on <em>reference sequences</em>. It outputs both a set of maximal unitigs and, critically, a <em>tiling</em>, that describes precisely how each reference is spelled out by an ordered sequence of unitigs (each in a specific orientation), and possibly with gaps of <code class="language-plaintext highlighter-rouge">N</code> nucleotides.  When we build a <code class="language-plaintext highlighter-rouge">piscem</code> index, we build an <code class="language-plaintext highlighter-rouge">SSHash</code> index over the unitig sequences, and a packed, inverted index over the tiling information. In this inverted index, we store, for each unitig, the sequence of references, positions and orientations in which it occurs.</p>

<p>However, when we (primarily <a href="https://sites.google.com/view/jamshed/home">Jamshed</a>) wrote <code class="language-plaintext highlighter-rouge">Cuttlefish</code> 1, we had in mind the construction over massive collections of <em>large</em> references (e.g. large bacterial pangenomes, plant genomes, or human genomes / pangenomes).  Specifically, we had in mind reference collections where, on average, each reference sequence was long (e.g. at least millions to tens of millions of nucleotides long).</p>

<p>This informed the parallelization strategy we adopted.  Each input reference was cut into a number of bins equal to the thread count, each thread wrote out the tiling information for its assigned interval to an intermediate file. Then, these intermediate tilings were stitched together in order to yield the tiling information for the entire reference sequence before we moved on to the next reference sequence (e.g. the next chromosome).  This worked brilliantly for genomes, but quite horribly for transcriptomes.  When each reference sequence is short (1-2 kilobases), each thread ends up being assigned a small interval of sequence (a few dozen to a few hundred bases long, and typically tens of unitigs), writing tiny intermediate tilings to file, forcing the flushing of tiny buffers, and then reading those all in again to stitch together the overall tiling.</p>

<p>All of these tiny I/Os just made this step horrifically slower than it needed to be, and this was <em>especially</em> a problem on networked file systems, where millions or billions of tiny flushed I/Os absolutely tanked performance.  An index build that should have taken 10s of minutes would end up taking many hours or, sometimes, even tens of hours.  There were quite a few GitHub issues about this, and even a friendly <a href="https://github.com/COMBINE-lab/simpleaf/issues/181">how-to guide</a> written by one of our users.</p>

<p>The solution was clear; avoid the intermediate files altogether. As the references are short, they aren’t needed.  Likewise, it’s silly to partition a reference into many bins between threads when each thread ends up writing only a few symbols in the tiling.  Instead, it makes <em>much</em> more sense to parallelize <em>across</em> reference sequences — to let each thread pull the next available reference sequence and create its tiling information in memory.</p>

<p>Conceptually, all of this was clear, but it was buried in a quite large C++ code base, and this specific code hadn’t been touched in several years. Jamshed and I discussed the fix, and it was on his radar, but he always had more important things to tackle, and then he graduated and moved on to his postdoc at Northeastern (where he’s doing awesome stuff, btw).</p>

<p>So, again, and based on my earlier success, I decided to give an AI coding agent a hand.  This time, I didn’t really want to have to paste bits and pieces of code into the Claude website. While I was curious about these agents, I still wasn’t curious enough to commit to paying real, hard, <em>money</em> for one.</p>

<p>Luckily, as part of our academic-affiliated organization (<code class="language-plaintext highlighter-rouge">COMBINE-lab</code>) on GitHub, we get the cheap tier of co-pilot (the $10 / month version). So, I hopped on the co-pilot interface on GitHub and started chatting with it to create the feature. There was a bit of back and forth conversation, but over the course of a few hours, co-pilot and I worked on <a href="https://github.com/COMBINE-lab/cuttlefish/pull/60">this PR</a> that implemented the feature.  This made the bottleneck step of <code class="language-plaintext highlighter-rouge">Cuttlefish</code> (on transcriptomes) <em>a lot</em> faster. It was about 5 times faster, even when running on a reasonably fast local disk, and much more than that when running on a cluster with a networked file system (though it <em>still</em> makes sense to do a build and write the output to local scratch and then move it to a shared location).  This experience was a <em>huge success</em>.  The AI models (mostly Claude Sonnet and Opus) “understood” the C++ codebase, my instructions, and our goal.  It went off and implemented the feature piece-by-piece, coming back to me when expert input was required.</p>

<p><em>Not everything was perfect</em>, however.  For example, when implementing the inter-reference parallelization strategy, we reached a point where the agent introduced a concurrency bug.  I manually investigated the code and immediately noticed the source of the error, and why it was happening. I pointed out where the error was, and let co-pilot take a shot. Once I pointed out where the problem was, it did come up with a solution, but it was a sub-optimal one. I explained my solution, and it agreed the solution I proposed was better.</p>

<p>Nonetheless, this was a minor hiccup in what was otherwise a big win.  Even when interacting with the agent in this way, a key ability that it had was to inspect the feedback I provided critically, and reason about how to address the relevant issues. At one point, when we encountered a segmentation fault, I popped open the executable in <code class="language-plaintext highlighter-rouge">lld</code>b, ran it, and provided the backtrace to co-pilot. Based on the backtrace, it reasoned about, diagnosed, and fixed the bug.</p>

<p>This experience led me to attempt a bigger lift.</p>

<h2 id="two-major-translations-in--1-week-actully-just-about-4-days">Two major translations in &lt; 1 week (actully just about 4 days)</h2>

<h3 id="sshash"><code class="language-plaintext highlighter-rouge">SSHash</code></h3>

<p><a href="https://github.com/jermp/sshash"><code class="language-plaintext highlighter-rouge">SSHash</code></a> is an incredibly efficient sequence dictionary for <code class="language-plaintext highlighter-rouge">k</code>-mer lookup in the context of genomics.  The original data structure is described in <a href="https://academic.oup.com/bioinformatics/article/38/Supplement_1/i185/6617506">this</a> paper by Giulio Ermanno Pibiri, the reading of which, actually, initially sparked my collaboration with Giulio (so thank goodness for this paper!).  Giulio and I <a href="https://www.biorxiv.org/content/10.64898/2026.01.21.700884v1">recently pre-printed</a> a set of data structural and algorithmic updates to <code class="language-plaintext highlighter-rouge">SSHash</code> aimed at theoretically and practically improving the caching behavior.</p>

<p>I’ve wanted a version of <code class="language-plaintext highlighter-rouge">sshash</code> in Rust now for a <em>long</em> time.  In fact, I asked my PhD student Jason Fan (now at Fulcrum Genomics) to try to implement it about 2 years ago!  Jason made some great progress, but it was a heavy lift, and not directly on the path of his dissertation, so this project was never finished.  Further, <code class="language-plaintext highlighter-rouge">sshash</code> <em>is</em> on the core path of our lightweight read mapping tool, <code class="language-plaintext highlighter-rouge">piscem</code>, that is, itself, used upstream of <code class="language-plaintext highlighter-rouge">alevin-fry</code> for single-cell RNA-seq and single-cell ATAC-seq processing, so it was a perfect candidate for a “large scale” project to convert from C++ to Rust.</p>

<p>I didn’t want to attempt this through the co-pilot web interface, so I downloaded the co-pilot CLI.  I was ready to go! And the co-pilot CLI … immediately segfaulted (another story for another post). So, I threw up my hands and decided to work through the VSCode co-pilot plugin.  It works well, but I don’t use VSCode, I’m a neovim guy. Anyway, I wasn’t to be doing most of the coding (just the chatting and guiding, with some minimal manual fixes), so I went with it.</p>

<p>What happened next was a <em>transformative</em> experience for me.  While converting an approximately single-file dependency for <code class="language-plaintext highlighter-rouge">mim</code> was very useful, and while directly modifying our <code class="language-plaintext highlighter-rouge">Cuttlefish</code> C++ codebase and adding a long-awaited feature was impressive, actually, what happened with <code class="language-plaintext highlighter-rouge">sshash</code> <em>massively</em> updated my prior about these tools.</p>

<p>I started with describing what <code class="language-plaintext highlighter-rouge">SSHash</code> is, giving the agent access to the full C++ <code class="language-plaintext highlighter-rouge">SSHash</code> codebase, and describing the design requirements for the Rust version. Claude derived a detailed implementation plan, broken down into major phases and minor sub-phases and steps within each phase.  It described what the requirements were for each phase, how correctness was to be assessed, and when it would be OK to move on to the next step.</p>

<p>Over the next approximately 2 days, in a few focused sessions, and sometimes in the background (between meetings, etc.), I guided Claude, through co-pilot, into a <em>complete</em>, <em>functional</em> and <em>approximately performance equivalent</em> version of <code class="language-plaintext highlighter-rouge">sshash</code> written in Rust.  Now, certainly, there was a lot of great (non-AI-created) infrastructure that we could rely on. <code class="language-plaintext highlighter-rouge">SSHash</code> is largely about succinct data structures, and we were able to build directly on <a href="https://github.com/vigna/sux-rs"><code class="language-plaintext highlighter-rouge">sux-rs</code></a> for bit-packed integer vectors and Elias Fano vectors, and on <a href="https://github.com/beling/bsuccinct-rs"><code class="language-plaintext highlighter-rouge">bsuccinct</code></a> for the incredible <a href="https://arxiv.org/abs/2504.17918"><code class="language-plaintext highlighter-rouge">PHast</code></a> minimal perfect hash function (it was a real win that so many of the leaders in the succinct data structure space are “Rust forward”).</p>

<p>Nonetheless, this implementation feat absolutely <em>blew me away</em>.  We went from 0 to a fully-functional and performance equivalent version of this very sophisticated index over 2 days; a goal that had eluded me for 2 years. Granted, if I was in a position to set aside a month or so for nothing but crack coding sessions, I think I could have accomplished this, but I don’t have a month or two to set aside like this.  The fact that, in such a short span, I now had access to <code class="language-plaintext highlighter-rouge">SSHash</code> in Rust was simply mind-blowing.</p>

<p>Again, as with the <code class="language-plaintext highlighter-rouge">Cuttlefish</code> feature. Not everything was perfect the first time. In one step, Claude replaced what was an efficiently encoded Elias Fano sequence with O(1) select support with a <code class="language-plaintext highlighter-rouge">Vec&lt;u64&gt;</code> with a binary search.  This was a silly and unnecessary mistake from both the size and speed perspective, and I have no idea <em>why</em> it did this. But, because I was watching this process in amazement, I noticed this, pointed out the issue, and it was promptly fixed.</p>

<p>Also, <em>very interestingly</em>, despite being given explicit and repeated instructions to directly follow the set of data structures used in the C++ implementation, Claude initially came up with a novel encoding scheme for offsets into light, medium, and heavy minimizer buckets, deriving an approach that was distinct from (but in practice competitive with) what was done in the C++ code. For the time being, I’ve left that design on a <a href="https://github.com/COMBINE-lab/sshash-rs/tree/initial-design">branch of the <code class="language-plaintext highlighter-rouge">sshash-rs</code> repository</a>, as I think it was a quite interesting development.</p>

<p>Yet, again, with some cajoling and oversight, Claude produced a semantically equivalent index to what exists in C++. This index now exists in the <a href="https://github.com/COMBINE-lab/sshash-rs/tree/main"><code class="language-plaintext highlighter-rouge">sshash-rs</code> repository</a>, and it was critical to me in my most recent project (yes, there was one after <code class="language-plaintext highlighter-rouge">sshash-rs</code> which was started and had the initial version completed this week).  I hope that having <code class="language-plaintext highlighter-rouge">SSHash</code> available in Rust also proves useful to other folks working in this space. It’s also astounding to me that all of this (modulo some subsequent optimizations for which I used Claude Code) was accomplished on a $10 / month co-pilot plan!</p>

<h3 id="piscem-rs-the-magnum-opus-nerd-pun-intended"><code class="language-plaintext highlighter-rouge">piscem-rs</code>, the magnum <em>Opus</em> (nerd pun intended)</h3>

<p>With the <code class="language-plaintext highlighter-rouge">sshash</code> Rust implementation in place, and enough evidence, finally, to pay for an actual Claude Max subscription. I booted it up, and started an ~1.5 day session that would ultimately result in the complete translation of our <a href="https://github.com/COMBINE-lab/piscem-cpp/tree/dev-pos"><code class="language-plaintext highlighter-rouge">piscem-cpp</code></a> mapping tool from C++ (with a submodule dependency on the C++ <code class="language-plaintext highlighter-rouge">SSHash</code>) to Rust (with a proper crate dependency on <code class="language-plaintext highlighter-rouge">sshash-rs</code>).</p>

<p>Like the <code class="language-plaintext highlighter-rouge">sshash</code> translation, this is a complex and interrelated codebase, implementing many features. <code class="language-plaintext highlighter-rouge">Piscem</code> supports mapping bulk RNA-seq data, single-cell (and single-nucleus) RNA-seq data, and single-cell ATAC-seq data, with an array of different algorithms and variants.  It supports output in a custom binary format (the <a href="https://hackmd.io/@PI7Og0l1ReeBZu_pjQGUQQ/HkbVOHXUR"><code class="language-plaintext highlighter-rouge">RAD</code></a> format), consumed downstream by our <a href="https://github.com/COMBINE-lab/alevin-fry"><code class="language-plaintext highlighter-rouge">alevin-fry</code></a> tool for single-cell quantification. It implements necessary supporting features, like parsing fragment geometries using a custom parsing expression grammar (PEG). It has a host of different detailed optimizations, like a concurrent shared cache to retain lookup information about the <code class="language-plaintext highlighter-rouge">k</code>-mers bridging the ends of unitigs and the unitigs that most frequently follow them in the observed reads (the unitig end cache from our paper <a href="https://academic.oup.com/bioinformatics/article/41/Supplement_1/i237/8199381">“Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment”</a>).</p>

<p>All of this was created, in about a day and a half, by me driving Claude Code (with Opus 4.6).  The result is astonishing.  The current version supports all of the features of the C++ codebase, is slightly <em>faster</em> than the corresponding C++ code, is more cleanly organized, and produces semantically-equivalent <code class="language-plaintext highlighter-rouge">RAD</code> format output (e.g. the order of things can be different because of multithreading). On non-trivial test data, across all assay types, there is <strong>100% concordance</strong> between the results of the C++ and Rust versions of <code class="language-plaintext highlighter-rouge">piscem</code>.</p>

<p>What’s even more amazing, perhaps, and something that truly sets the experience of using Claude Code apart from something like the web interfaces or even co-pilot in VSCode, is Claude Code’s tremendous ability to interact with the system, plan what it is doing, and <em>introspect</em> on the results of the code it is generating. During this process, when there were performance pitfalls or subtle differences in the algorithm. Claude Code instrumented the code with debug information and inspected the output to understand what was going wrong. <em>Claude Code</em> wrote the test harness to compare the binary <code class="language-plaintext highlighter-rouge">RAD</code> format files for <em>semantic</em> equivalence between the C++ and Rust versions of <code class="language-plaintext highlighter-rouge">piscem</code>. <em>Claude Code</em> wrote the unit tests to ensure that each abstraction worked in isolation and that implementation changes didn’t affect the output. <em>Claude Code</em> instrumented the index build on an ATAC-seq index (which is larger than that used for transcriptome mapping, because it indexes the whole genome) when I complained that it took too long, discovered the performance issue arising from a pathologically cache incoherent memory access pattern, fixed the pattern in the dependent codebase, validated the performance and correctness of this fix, and then asked me if I wanted to commit and push this change (I did).</p>

<p>In the series of translating the C++ code to Rust, we got down to a case where there was <em>one particular</em> sequencing read being mapped by the C++ program that was not being mapped by the Rust program.  Claude suggested this was a very minor difference, but I insisted we fix it.  Claude isolated the read, tracked down the difference, and explained why the read mapped in the C++ code and not the Rust code. It turns out, it was the result of a (very rare) integer overflow case in masking the bitpacked representation of a length 31 k-mer in a 64-bit integer.  Claude pointed out what this bug in the C++ code was, why it happened, and offered to reproduce the buggy behavior in Rust to achieve perfect parity.  Of course, I chose to simply fix the C++ code instead, but just going on that journey with Claude Code, and seeing it track down a rare corner case and reason about the algorithmic flow, was eye-opening.  A huge part of what makes Claude Code so powerful, in my opinion, is the way that it can (and often expertly does) <strong>close the loop</strong> on developing, debugging and optimizing.</p>

<p>As a result of this development push, in about 2 days, we now have a complete <a href="https://github.com/COMBINE-lab/piscem-rs">implementation of <code class="language-plaintext highlighter-rouge">piscem</code> in Rust</a>.  This will be shortly replacing our C++ implementation, making it easier to add new features, making it easier for my students to contribute, and making maintenance and deployment much simpler!</p>

<p><strong>In the future</strong>, I hope to apply these tools to both novel development, as well as some further translations to Rust that I’ve long desired (I’m looking at you, <a href="https://github.com/COMBINE-lab/salmon"><code class="language-plaintext highlighter-rouge">salmon</code></a>).  Also, along with a team of interested individuals, we have a super-secret (well, maybe not so secret if you follow me on BlueSky) Claude assisted rewrite planned, and I’m super psyched about this.</p>

<h3 id="non-technical-caveats">Non-technical caveats</h3>

<p>I’ve substantially revised my perspective on the capabilities and utility of these AI coding models, and I think that they are and will continue to be an incredibly powerful tools, and I am excited to continue to explore how they can help achieve technical (and specifically software) goals. Nonetheless, there are several very real, and very serious caveats to these AI models in general that I think are quite concerning, and for which I don’t currently have a solution.  I mention a few below. However, I <em>do not</em> think that simply refusing to engage with these tools is a productive way to address these caveats. I do not think that, even if we wanted, the genie could be put back in the bottle.</p>

<p>Here’s a super-short (and non-exhaustive) list of issues with these tools, how they have come to be, and the effects that will have that keep me up at night. I may likely continue to update this, or even turn it into its own post.</p>

<ul>
  <li>
    <p>Training by violating copyright and replicating intellectual property with zero traceability or citation. Despite the legal rulings on fair use exemptions, it seems to me quite clear that these models, in their training, have massively violated relevant copyright and licensing terms. They have consumed, with almost complete disregard for any protection, the intellectual property of countless individuals. Further, by virtue of the way these models work, at least as far as I am aware, when they recapitulate or even completely reproduce this subsumed IP, they provide no citation nor attribution.  I do not know how to hold accountable the companies who have built these models for engaging in practices where they <em>certainly</em> know better, or even if that is possible. But I do hope, that going forward, we can develop both social <em>and</em> technological solutions to this serious problem.</p>
  </li>
  <li>
    <p>Brainrot: My experience with these models has given me a bit of a “black mirror” moment. What the models do, the capabilities they have, and the quality of the content they produce is, to a large extent, a reflection of the user.  They can be an incredible tool for a serious an seasoned developer building software, or for a mathematician working through a complex open conjecture, and even for an ambitious and dedicated teenager looking to master a new skill. However, they can also, <em>absolutely</em>, act as an “easy” button to generate zero-to-little effort content; output that may be passable, but is of no real value. It seems clear that we are well into the regime where a student can choose to put in almost no effort, have one of these models complete a (e.g. programming) assignment for them, and completely miss out on the educational goal.  While there is a lot of thought going into how to incorporate AI tools <em>into</em> curricula, and while that may be an important line of inquiry, I think it is incomplete.  The real problem that we need to tackle, that predates the challenges raised by AI, but that is newly-magnified by their capabilities, is the very human and social one. How do we motivate actual learning? How do we convince people that learning, itself, is the goal, and to put in real effort when there may be easy ways to game the evaluation system (i.e. grades)? Motivated and truly curious people will learn with or without AI tools, but these tools make it much easier for folks not so-inclined to move through the educational system without learning or exercising critical thinking skills, and perhaps even fooling themselves into thinking they are.</p>
  </li>
  <li>
    <p>Widening the gap: The more I experiment with these tools and the more I see how capable they are, the more I also see how savant like they can be. Their capabilities are incredible, but their mistakes are often <em>absolutely silly</em> (<strong>to a seasoned expert</strong>).  Building on the “black mirror” comment above, I think that it may be the case that, rather than having a “rising tide” that “raises all ships”, these tools may further widen the divide between the most effective, capable and knowledgeable experts, and everyone else.  It is the seasoned greybeard, who has spent months of cumulative time tracking down and squashing subtle heisenbugs, who can immediately see and correct the silly but obvious (to them) mistake that Claude just made.  It is the accomplished algorithm engineer who knows exactly how a specific data structure needs to be laid out to maximize cache efficiency and therefore performance. It is that algorithm engineer, and the knowledge they bring to the tool, that lets them guide the model to the <strong>right</strong> solution, and not just <strong>a</strong> solution.  Ultimately, I fear that if we are not able to effectively teach and instill that knowledge and experience in those who are now undergoing that critical stage of their development, we may be creating, in some ways, an expertise cliff. For those who have that expertise today, it has often come through hard-earned knowledge, manual construction of sophisticated systems from first principles, and a <em>lot</em> of persistence and banging their heads against the wall. There may be ways to create that level of expertise and knowledge without all of the associated “manual” exercise, but those ways are as yet, to me, unclear.</p>
  </li>
</ul>]]></content><author><name>Rob Patro</name></author><category term="blog" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">Why use Rust for bioinformatics? Part 2: You can depend on me.</title><link href="https://arte-lab.github.io/blog/2022/11/28/rust-for-bioinformatics-part-2.html" rel="alternate" type="text/html" title="Why use Rust for bioinformatics? Part 2: You can depend on me." /><published>2022-11-28T00:00:00+00:00</published><updated>2022-11-28T00:00:00+00:00</updated><id>https://arte-lab.github.io/blog/2022/11/28/rust-for-bioinformatics-part-2</id><content type="html" xml:base="https://arte-lab.github.io/blog/2022/11/28/rust-for-bioinformatics-part-2.html"><![CDATA[<p>For part 2 in our “Why use Rust in bioinformatics?” series, I want to focus on one of my favorite parts of the Rust ecosystem, <a href="https://doc.rust-lang.org/stable/cargo/">Cargo</a>. In fact, there is so much to like about Cargo, that I won’t even be able to cover that in a single post. Instead, I’ll focus in this post mostly on the use of Cargo for dependency management and will likely return later to some of my favorite <code class="language-plaintext highlighter-rouge">cargo</code> commands / plugins (like <a href="https://github.com/rust-lang/rust-clippy"><code class="language-plaintext highlighter-rouge">clippy</code></a>).</p>

<h4 id="what-is-cargo">What is Cargo?</h4>

<p>Cargo is the package manager for Rust and, more than that, it’s essentially the build system and <em>project management system</em>.  Cargo can be used to <a href="https://doc.rust-lang.org/cargo/commands/cargo-init.html">initialize the skeleton for a new project</a>, to <a href="https://doc.rust-lang.org/cargo/commands/build-commands.html">build your program’s executables or libraries</a>, to <a href="https://doc.rust-lang.org/cargo/commands/cargo-test.html">run your unit or integration tests</a>, to <a href="https://doc.rust-lang.org/cargo/commands/cargo-rustdoc.html">generate the documentation for your program</a>, to <a href="https://doc.rust-lang.org/cargo/commands/cargo-bench.html">run benchmarks</a>, and to perform a host of other useful actions.</p>

<p>In fact, Cargo does so much that I’m not going to attempt to cover it all in this post.  There is entire <a href="https://doc.rust-lang.org/cargo/index.html">online book</a> dedicated to Cargo, its use, and its capabilities.  Instead, I’m going to focus mostly on Cargo’s function as a depndency / package manager.</p>

<p>So, before I go into details, the <strong>TLDR</strong> is that Cargo is an amazing, easy-to-use, powerful, and intuitive build system that makes pulling in dependencies trivial, provides mechanisms for semantic versioning-based dependency resolution, reproducible builds, and automatic upgrading of dependency versions.  More than build systems I’ve encountered in any other language, Cargo “just works”, and it makes building projects in Rust, even those with substantial sets of dependencies, quick and easy.</p>

<h4 id="declaring-dependencies-with-cargo">Declaring dependencies with Cargo</h4>

<p>Cargo relies on a <a href="https://toml.io/en/">TOML</a> format file called <code class="language-plaintext highlighter-rouge">Cargo.toml</code> that describes certain metadata about your project, including its developers, what it does, how it is structured, its dependencies and its relevant compiler options.  At a high level, Cargo is a declarative system (it is possible to construct “imperative” build scripts — so-called <code class="language-plaintext highlighter-rouge">build.rs</code> files — but they are not needed for many projects), and <em>building your project is as simple as declaring what type of project it is, listing your dependencies and preferred compiler options, and running <code class="language-plaintext highlighter-rouge">cargo build --release</code></em>.</p>

<p>As a non-trivial running example, I’ll be using the <code class="language-plaintext highlighter-rouge">Cargo.toml</code> file from our <a href="https://github.com/COMBINE-lab/alevin-fry/"><code class="language-plaintext highlighter-rouge">alevin-fry</code></a> tool for single-cell and single-nucleus RNA-seq processing.  The first part of the file describes the package, including metadata like the package name, version, authors, etc.  Now, not all of these fields are strictly required, but it’s nice to populate your <code class="language-plaintext highlighter-rouge">Cargo.toml</code> with relevant metadata that will make tracking and organizing it easier in the context of other packages.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[package]
name = "alevin-fry"
version = "0.8.0"
authors = [
  "Avi Srivastava &lt;avi.srivastava@nyu.edu&gt;",
  "Hirak Sarkar &lt;hirak_sarkar@hms.harvard.edu&gt;",
  "Dongze He &lt;dhe17@umd.edu&gt;",
  "Mohsen Zakeri &lt;mzakeri@cs.umd.edu&gt;",
  "Rob Patro &lt;rob@cs.umd.edu&gt;",
]
edition = "2021"
description = "A suite of tools for the rapid, accurate and memory-frugal processing single-cell and single-nucleus sequencing data."
license-file = "LICENSE"
readme = "README.md"
repository = "https://github.com/COMBINE-lab/alevin-fry"
homepage = "https://github.com/COMBINE-lab/alevin-fry"
documentation = "https://alevin-fry.readthedocs.io/en/latest/"
include = [
  "/libradicl/src/*.rs",
  "/src/*.rs",
  "/Cargo.toml",
  "/README.md",
  "/LICENSE",
  "/CONTRIBUTING.md",
  "/CODE_OF_CONDUCT.md",
]
keywords = [
  "single-cell",
  "preprocessing",
  "RNA-seq",
  "single-nucleus",
  "RNA-velocity",
]
categories = ["command-line-utilities", "science"]
</code></pre></div></div>

<p>Most of these fields are self-explanatory, and the syntax is quite straightforward. The entries are a series of key-value pairs, where the value can be a string, a list, or (as we’ll see below) a nested key-value store.  Perhaps the only non self-explanatory key here is the <code class="language-plaintext highlighter-rouge">edition</code> key.  The idea of rust <code class="language-plaintext highlighter-rouge">edition</code>s are described <a href="https://doc.rust-lang.org/edition-guide/editions/index.html">here</a>, and they essentially describe small backwards incompatible language changes as well as certain default behaviors.  Currently “2021” is the most recent <code class="language-plaintext highlighter-rouge">edition</code> of rust, and that is what we set here.</p>

<p>The actual dependencies are declared in a section labeled — unexpectedly — as “dependencies”.  A short excerpt is below:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[dependencies]
# for local development, look in the libradicl git repository
# but when published, pull the specified version
libradicl = { git = "https://github.com/COMBINE-lab/libradicl", version = "0.4.6" }
anyhow = "1.0.59"
arrayvec = "0.7.2"
ahash = "0.7.6"
bincode = "1.3.3"
bstr = "0.2.17"
</code></pre></div></div>

<p>This demonstrates several important details about the dependency management system exposed by Cargo.  The first thing is the simple manner in which dependencies are declared.  Each dependency is the name of a <em>crate</em> (the terminology that <code class="language-plaintext highlighter-rouge">Cargo</code> uses for dependencies), followed by a version constraint.  In general, Cargo crates follow semantic versioning, and the default syntax for declaring a compatible versions “X.Y.Z” means that you are willing to accept any version <em>compatible</em> with “X.Y.Z”.  So this would match, for example, “X.Y.(Z+1)” or “X.Y.(Z+2)”, but not “X.(Y+1).Z”.  You can also declare a constraint as “X.Y” which would allow “X.(Y+1).Z” but not “(X+1).Y.Z” etc. You can even declare dependencies as “X”, which allows any version &gt;=X and &lt;(X+1).  The full syntax for specifying dependency constraints is quite powerful and flexible, and you can read more about it <a href="https://doc.rust-lang.org/cargo/reference/specifying-dependencies.html">here</a>.</p>

<p>The second thing to note is that some dependencies have a more complex description.  For example, the first dependency is <code class="language-plaintext highlighter-rouge">libradicl</code>, a library that we also developed that is hosted on GitHub as well as on <code class="language-plaintext highlighter-rouge">crates.io</code>.  You’ll note that the declaration of the dependency lists two different sources, a <code class="language-plaintext highlighter-rouge">git</code> source and a <code class="language-plaintext highlighter-rouge">version</code> source.  This is a nice feature of Cargo.  When the program is built locally, it will pull the relevant dependency from the GitHub repository listed in the URL.  This allows one to develop coupled packages with ease, by allowing a program to always pull in its dependency with the latest commits directly from a remote (or local) repository.  Yet, when your package is later <em>published</em> (more on that when we talk about <code class="language-plaintext highlighter-rouge">crates.io</code>), it can’t rely on dependencies tracked in Git.  For that, you must instead declare a dependency on another crate that is hosted on <code class="language-plaintext highlighter-rouge">crates.io</code> — here, we rely on version 0.4.6 of the <code class="language-plaintext highlighter-rouge">libradicl</code> crate (or any version compatible with this declaration).  This leads to a fairly fluid development experience, where, when working on the <code class="language-plaintext highlighter-rouge">alevin-fry</code> tool, if we need to make a change to <code class="language-plaintext highlighter-rouge">libradicl</code>, we first make the changes upstream in GitHub (pulling directly from the repo during development).  Then, when we are satisfied with the changes that we wish to make, we push a new version to <code class="language-plaintext highlighter-rouge">crates.io</code>, and bump the <code class="language-plaintext highlighter-rouge">version</code> string in the <code class="language-plaintext highlighter-rouge">libradicl</code> dependency to this new version.  It’s also worth noting the ease with which the <a href="https://github.com/google-github-actions/release-please-action"><code class="language-plaintext highlighter-rouge">release-please</code> GitHub action</a> and <a href="https://github.com/actions-rs/cargo"><code class="language-plaintext highlighter-rouge">rust</code> action</a> allows tagging a new version and automatically pushing the resulting release to <code class="language-plaintext highlighter-rouge">crates.io</code>.</p>

<p>If you look farther down in the <code class="language-plaintext highlighter-rouge">Cargo.toml</code> file, you’ll notice that some other dependencies contained different invocations in their declarations. While the documentation provides a full accounting of how these different properties work, most of them are actually rather self-explanatory.  For example, the declaration below is a dependency on the brilliant <code class="language-plaintext highlighter-rouge">serde</code> serialization crate.  In addition to the version constraint, we also have a property <code class="language-plaintext highlighter-rouge">features = ["derive"]</code>. In rust, crates may have default and optional “features”, these describe functionality that the crate can be built without or that it can provide.  Here, we are declaring that we wish to enable the “derive” feature of the <code class="language-plaintext highlighter-rouge">serde</code> crate (which lets us use the <code class="language-plaintext highlighter-rouge">derive</code> macro to quickly build out serialization and deserialization for the structs and types in our program).</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>serde = { version = "1.0.136", features = ["derive"] }
</code></pre></div></div>

<h4 id="declaring-dependency-version-constraints">Declaring dependency version constraints</h4>

<p>Cargo allows several ways to declare constraints on dependency versions.  The default behavior “X.Y.Z” is equivalent to the syntax “~X.Y.Z” which restricts the dependency from being satisfied by another version that makes breaking changes.  If you want or need to pin your dependency to a <em>specific</em> version, you can use the syntax “=X.Y.Z”, which will require pulling down exactly this version of the corresponding crate. There are many other types of constraints you can place on the dependencies (e.g. “&gt;X.Y.Z”, etc.).  These various constraints and how they work is documented nicely in <a href="https://doc.rust-lang.org/stable/cargo/reference/specifying-dependencies.html">the book</a>.</p>

<h4 id="resolving-dependencies-and-the-cargolock-file">Resolving dependencies (and the Cargo.lock file)</h4>

<p>When you ask Cargo to build your program, it will resolve the relevant dependencies, downloading the corresponding crates from <code class="language-plaintext highlighter-rouge">crates.io</code> (or other sources like GitHub if you have specified those) and building them to be linked with your program.  In the process of doing so, it’s performing dependency resolution.  That is, it will find a corresponding set of versions that, given the current state of <code class="language-plaintext highlighter-rouge">crates.io</code> (i.e. the current set of available versions of all of the crates on which your program depends), satisfies all of the constraints on versions you requested. Generally, subject to these constraints, it pulls down the newest possible versions.  This behavior is great, because this means that if a corresponding crate updates their latest available version with something that is compatible (in terms of semantic versioning and your specified constrains), then you can get this updated version just by re-building your program.</p>

<p>Of course, there are situations where, for the purposes of reproducible builds, you may wish to be a bit more strict on how dependencies are resolved.  Cargo’s way of allowing this is what is called the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file.  The contents of the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file look somewhat different than those of the <code class="language-plaintext highlighter-rouge">Cargo.toml</code> file (and they are generated by Cargo itself, so you’re not responsible for making this), and <a href="https://doc.rust-lang.org/cargo/guide/cargo-toml-vs-cargo-lock.html">the book has a section on these</a>. For example, an entry from the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file for <code class="language-plaintext highlighter-rouge">alevin-fry</code> looks like this:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[[package]]
name = "anyhow"
version = "1.0.65"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "98161a4e3e2184da77bb14f02184cdd111e83bbbcc9979dfee3c44b9a85f5602"
</code></pre></div></div>

<p>This specifies the <em>specific</em> resolution for a dependency that was obtained during the solve when the program was built. When Cargo attempts to build your program, before it attempts to check the available upstream crate versions and resolve your program’s dependencies, it first checks for the existence of a <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file.  <strong>If this file is present</strong>, then it will simply use the versions declared therein (that is, it will re-use this “solve” of your set of dependencies).  One thing that’s really nice about this is that it’s possible to essentially “freeze” a build using the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file, such that, if some upstream dependency fails to properly use semantic versioning and makes a breaking change with a “patch” bump, builds that use the successful <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file won’t be affected.</p>

<p>The standard recommended practices around <code class="language-plaintext highlighter-rouge">Cargo.lock</code> files is actually something that I only learned relatively recently. Initially, I’d assumed that these generated files were essentially not for user consumption, and so I added the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> files to my <code class="language-plaintext highlighter-rouge">.gitignore</code> list for my repositories and kept them out of version control (they are small, so this was for the purposes of keeping a clean file history rather than for worrying about repository size).  However, I since learned that recommended practice is basically the following: <em>If you are building a user-facing program or tool</em>, then you should include the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file in your version control and in your set of distributed source files; <em>On the other hand, if you are building a library</em> for other tools to pull in and depend upon, then you should generally not include the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file in your version control and distributed source files. Nonetheless, the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file is a neat solution to the problem of reproducible builds and solves.  Even when a specific dependency is “yanked” from <code class="language-plaintext highlighter-rouge">crates.io</code> (essentially, the authors of a crate can “unlist” a specific version of their crate), if you are in possession of the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file that solved using this yanked crate, your build will still be able to pull it down and compile against it.  In other words, even if certain versions of a crate are no longer publicly listed, the <code class="language-plaintext highlighter-rouge">Cargo.lock</code> file lets you re-create a build with the precise versions used before, making it easy to reproduce the set of dependencies of a previous build exactly.</p>

<h4 id="cargo-edit">Cargo-edit</h4>

<p>Cargo has a <em>plethora</em> of different commands it exposes, many are built in and some come a “plugins” that expand the capabilities of Cargo. One particularly cool plugin that I wanted to mention is <a href="https://crates.io/crates/cargo-edit"><code class="language-plaintext highlighter-rouge">cargo-edit</code></a>, and specifically, the <code class="language-plaintext highlighter-rouge">cargo upgrade</code> command. When your project has several dependencies, tracking and upgrading those dependencies can be a pain.  The <code class="language-plaintext highlighter-rouge">cargo-edit</code> plugin provides commands to help manage the contents of your <code class="language-plaintext highlighter-rouge">Cargo.toml</code> file, including the <code class="language-plaintext highlighter-rouge">add</code> command to add an entry for a new crate (given its name and the set of features you want), and to remove (<code class="language-plaintext highlighter-rouge">rm</code>) dependencies. It also includes a brilliant <code class="language-plaintext highlighter-rouge">upgrade</code> command that scans your list of dependencies, determines which can be safely upgraded, modifies your <code class="language-plaintext highlighter-rouge">Cargo.toml</code> accordingly, and also reports which packages can’t be upgraded given your current versions and constraints. For example, running <code class="language-plaintext highlighter-rouge">cargo upgrade</code> on <code class="language-plaintext highlighter-rouge">alevin-fry</code> (as of commit <code class="language-plaintext highlighter-rouge">a77c96e162758e8cf5f4e509263216158bb580c9</code>) gives the following output:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>    Updating 'https://github.com/rust-lang/crates.io-index' index
    Checking alevin-fry's dependencies
name            old req compatible latest  new req note
====            ======= ========== ======  ======= ====
ahash           0.8.1   0.8.2      0.8.2   0.8.2
crossbeam-queue 0.3.6   0.3.8      0.3.8   0.3.8
flate2          1.0.24  1.0.25     1.0.25  1.0.25
serde           1.0.147 1.0.148    1.0.148 1.0.148
serde_json      1.0.87  1.0.89     1.0.89  1.0.89
snap            1.0.5   1.1.0      1.1.0   1.1.0
chrono          0.4.22  0.4.23     0.4.23  0.4.23
mimalloc        0.1.31  0.1.32     0.1.32  0.1.32
clap            =3.2.16 3.2.16     4.0.27  =3.2.16 pinned
   Upgrading recursive dependencies
note: Re-run with `--pinned` to upgrade pinned version requirements
note: Re-run with `--verbose` to show all dependencies
  unchanged: anyhow, arrayvec, bincode, bio-types, bstr, crossbeam-channel, csv, indicatif, itertools, libradicl, needletail, num-format, num_cpus, petgraph, rand, rust-htslib, sce, scroll, slog, slog-async, slog-term, smallvec, sprs, statrs, thiserror, typed-builder
</code></pre></div></div>

<p>So we see what the current upgradable dependency is, the latest compatible version, the latest version (ignoring compatibility), and version to which our dependency has been upgraded. Finally, as you can see in the case of the <code class="language-plaintext highlighter-rouge">clap</code> dependency, if you have specific constraints that preclude upgrading a crate, it will also include relevant notes. While the <code class="language-plaintext highlighter-rouge">upgrade</code> command will not perform version breaking upgrades by default, you can pass the <code class="language-plaintext highlighter-rouge">-i, --incompatible</code> option to allow upgrading to an incompatible version and the <code class="language-plaintext highlighter-rouge">-p, --package</code> argument to target a specific package.  The command also has a <code class="language-plaintext highlighter-rouge">--dry-run</code> flag to show you what upgrades would be made without actually performing them.</p>

<p>Overall, <code class="language-plaintext highlighter-rouge">cargo-edit</code> makes adding, removing, and upgrading your dependencies easy, by taking the monotonous grunt work out of parts of the process that really should be automated.</p>

<h4 id="cratesio--the-source-for-official-dependencies"><code class="language-plaintext highlighter-rouge">crates.io</code> — The source for official dependencies</h4>

<p>I have mentioned <a href="https://crates.io/"><code class="language-plaintext highlighter-rouge">crates.io</code></a> above many times.  It is the official registry for rust language crates (dependencies), and currently home to &gt;98,000 different crates!  You can browse the crates by category or search for them by name. Moreover, when you start building your own libraries and tools, you can easily host them on <code class="language-plaintext highlighter-rouge">crates.io</code> for free. All you have to do is register and use the <code class="language-plaintext highlighter-rouge">cargo publish</code> command to upload your locally developed crate to the <code class="language-plaintext highlighter-rouge">crates.io</code> registry. After that, you (and others) can add dependencies on your crate simply by adding the appropriate declaration to your <code class="language-plaintext highlighter-rouge">Cargo.toml</code> file as we have discussed above.  In my opinion, one of the brilliant things about the rust ecosystem, is how the ease of both using <em>and publishing</em> crates encourages the development of small, modular, and reusable components in rust software.  There are crates for a host of different purposes, and it’s trivial to make your own. When you make a crate to serve a specific purpose, it is then easy to reuse it across many projects by simply declaring it as a dependency.  In my opinion, this works <em>much</em> better than the alternatives in languages like C/C++, where it is common to either vendor your dependencies and copy (potentially different versions) into the source tree across many projects that use them.  While certain package management solutions for C++ exist, like <a href="https://conan.io/"><code class="language-plaintext highlighter-rouge">conan</code></a> and <a href="https://vcpkg.io/en/index.html"><code class="language-plaintext highlighter-rouge">vcpkg</code></a>, these are all 3rd party solutions and they lack the scope and breadth of <code class="language-plaintext highlighter-rouge">cratres.io</code>, and also the tight and elegant integration with the rest of the development ecosystem that is enjoyed by <code class="language-plaintext highlighter-rouge">crates.io</code> and <code class="language-plaintext highlighter-rouge">cargo</code>. In my (admittedly biased) opinion, the dependency management solutions provided by rust are phenomenal, and probably the best among any language in which I’ve worked — and this includes non-compiled languages such as Python and R.</p>

<h4 id="some-fun-bioinformatics-crates">Some fun bioinformatics crates</h4>

<p>I’ll close this post by mentioning that a search for bioinformatics on <code class="language-plaintext highlighter-rouge">crates.io</code> turn up <a href="https://crates.io/search?q=bioinformatics">156 results</a> (and related terms turn up more e.g. <a href="https://crates.io/search?q=genomics">genomics turns up 112</a>).  I encourage you to go exploring yourself!  However, it is worth mentioning some common crates in the bioinformatics space that are pretty awesome:</p>

<ul>
  <li>
    <p>The <a href="https://crates.io/crates/bio"><code class="language-plaintext highlighter-rouge">bio</code> crate</a> is a bioinformatics library for Rust that provides implementations of several critical data structures (e.g. the FM-index) algorithms (e.g. alignment) and parsers (e.g. GTF). It’s a great place to start if you’re looking for a crate that tackles many common problems</p>
  </li>
  <li>
    <p>The <a href="https://crates.io/crates/seq_io"><code class="language-plaintext highlighter-rouge">seq_io</code> crate</a> is a particularly fast FASTA/Q parser. There are several such crates, so it’s worth exploring the different options here.</p>
  </li>
  <li>
    <p>The <a href="https://crates.io/crates/rust-htslib"><code class="language-plaintext highlighter-rouge">rust-htslib</code> crate</a> provides Rust bindings for the venerable <a href="https://github.com/samtools/htslib"><code class="language-plaintext highlighter-rouge">htslib</code></a> C library for reading and writing SAM/BAM/CRAM files.</p>
  </li>
  <li>
    <p>Related to the above, the <a href="https://crates.io/crates/noodles"><code class="language-plaintext highlighter-rouge">noodles</code></a> crate provides readers and writers for “BAM 1.6, BCF 2.2, BED, BGZF, CRAM 3.0, CSI, FASTA, FASTQ, GFF3, GTF 2.2, SAM 1.6, tabix, and VCF 4.3.” written entirely in Rust (so no binding to an external C library). It’s definitely a crate to keep an eye on in terms of native Rust support for these common file formats.</p>
  </li>
  <li>
    <p>The <a href="https://crates.io/crates/debruijn"><code class="language-plaintext highlighter-rouge">debruijn</code></a> crate from 10x genomics provides a de Bruijn graph implementation in Rust. In fact, 10x is quite prolific in terms of creating Rust crates in the bioinformatics space, many of which you can find <a href="https://crates.io/teams/github:10xgenomics:crates_io">here</a> — including a <a href="https://crates.io/crates/boomphf">Rust implementation of the BBhash minimal perfect hashing algorithm</a>.</p>
  </li>
  <li>
    <p>If you’re doing sequence alignment in Rust, definitely check out the <a href="https://crates.io/crates/block-aligner"><code class="language-plaintext highlighter-rouge">block-aligner</code></a> crate by <a href="https://twitter.com/daniel_c0deb0t">Daniel Liu</a>, for a high-performance, SIMD-accelerated, block-adaptive sequence alignment algorithm.</p>
  </li>
</ul>

<p>This is in no way a comprehensive list, but I would <em>absolutely</em> appreciate feedback if there are crates that you use regularly that you’d like me to list here! There are already a ton of great <em>tools</em> in Rust in the bioinformatics space, but the list above is mostly for library-level / reusable components.</p>]]></content><author><name>Rob Patro</name></author><category term="blog" /><category term="rust" /><category term="programming" /><category term="bioinformatics" /><category term="tools" /><category term="computer science" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry><entry><title type="html">Why use Rust for bioinformatics? Part 1: Defining the problem space.</title><link href="https://arte-lab.github.io/blog/2022/11/25/rust-for-bioinformatics-part-1.html" rel="alternate" type="text/html" title="Why use Rust for bioinformatics? Part 1: Defining the problem space." /><published>2022-11-25T00:00:00+00:00</published><updated>2022-11-25T00:00:00+00:00</updated><id>https://arte-lab.github.io/blog/2022/11/25/rust-for-bioinformatics-part-1</id><content type="html" xml:base="https://arte-lab.github.io/blog/2022/11/25/rust-for-bioinformatics-part-1.html"><![CDATA[<p>As has been well noted on the interwebs, I am a staunch advocate of the <a href="https://www.rust-lang.org">Rust</a> 
programming language.  This is particularly true in my home domain of bioinformatics and computational biology.
In fact, so persistent am I in my advocacy for the use of Rust in bioinformatics applications that <a href="https://kamimrcht.github.io/webpage/">some of my colleagues</a>
have claimed <a href="https://twitter.com/CamilleMrcht/status/1522609312006344705?s=20&amp;t=xH6INShi5gwSSHIZlqfoDw">I am an overfit bot</a>.</p>

<p>So, an obvious question a reader may ask is “why?”.  Why am I so zealous in my advocacy for Rust and its use in 
Bioinformatics?  What benefits do I think it provides over the alternatives? What even are the alternatives? 
Are there places in bioinformatics where I <em>don’t</em> think Rust is the right choice?</p>

<p>I intended this post to be an initial foray into addressing these questions, but quickly realized that, at least with 
the current queue of things that need my attention, a sprawling, long-form blog post was not the way to go.
Therefore, I am hopeful that this will become a series of posts, each one small and bite-sized, that together 
lay out some of my thoughts on the use of Rust in bioinformatics, each touching on a smaller piece of the whole 
picture. Yet, my prior blogging discipline is lacking, so I will make promises at this point about the length or 
frequency of this series.  So, without further ado, let’s get started.</p>

<p>In today’s post, I simply wish to define the problem space, that is usually implicit in my comments, when I advocate
for the use of Rust in bioinformatics.</p>

<h3 id="defining-the-problem-space">Defining the problem space</h3>

<p>When I advocate for the use of Rust as a language for developing bioinformatics methods and tools, I do so from my own
particular perspective. Bioinformatics is a <em>giant</em> field, with many sub-disciplines, problem areas, and methodological
approaches.  Perhaps, Rust is applicable everywhere here, but that is not the argument I mean to make.  Rather, I would like 
to advocate for the use of Rust when developing tools and methods for <em>data-intensive</em>, <em>high-throughput</em> analysis.</p>

<p>The types of applications I have in mind are sequencing indexing, read mapping and alignment, genome and transcriptome assembly, bulk and
single-cell RNA-seq and metagenomic quantification, etc.  These applications are characterized by a need to process a large volume 
of input data — often what many consider as <em>raw</em> input data. These applications have several characteristics that tie them together.
They often require reading and sometimes writing large quantities of data.  They often require low-level and / or binary parsing and 
interpretation of records.  Given the problem sizes we encounter in practice, these are usually applications where memory usage is a 
concern and so they often rely on efficient (in space and time) implementation specific data structures.  Further, in many such applications,
memory needs are often semi-regular and predictable.  Also, these applications often have components or subroutines that 
can be made embarrassingly parallel.  There are other things such problems have in common, but these will suffice for now.</p>

<p>This, of course, leaves out huge areas of bioinformatics — problems where we have processed or pre-processed data, and we want to perform 
exploratory data analysis, or particular types of statistical testing, or do certain types of dynamical modeling.  It is not that I <em>don’t</em>
think Rust could be a compelling choice for these types of applications (though I have my doubts about exploratory data analysis), it is just 
that these are not the application areas where I typically work, and so they are not the areas where I have the experience or confidence 
to advocate for Rust (yet).</p>

<h3 id="so-what-other-languages-occupy-this-space">So what other languages occupy this space?</h3>

<p>To argue for what makes Rust compelling, I first have to lay out what I think are the other common (and, perhaps, not-so-common) languages
used for these types of problems.</p>

<h4 id="c-and-c">C and C++</h4>

<p>Perhaps the most common languages are C and C++.  I mention these together, as is common, but it is 
critical to understand that they are <em>very</em> different languages.  The C language is relatively small and in many senses minimalistic. It
provides a small set of tools for modeling problems and for interacting with the system. It has a standard library, but a comparatively 
small one by modern standards, and things like collections must be written or provided as third-party libraries. I won’t say too much 
else about C here specifically, except that almost all the issues I raise about safety with respect to C++, and often times they are
even worse.  Since the C type system isn’t as rich as the one in C++, C often makes specific procedures generic by simply casting around 
data into a payload (e.g. a <code class="language-plaintext highlighter-rouge">void*</code>) rather than generating type-safe code for specific invocations at compile-time (e.g. C++’s templates 
and monomorphization).</p>

<p>As opposed to C, the C++ language is a giant monster.  First, you must ask, to which C++ is one referring? C++98, C++03, C++11, C++14, C++17, C++20?
If we consider the modern variants of C++ (e.g. C++11 and later), then these languages have added many useful features (but also a huge amount of 
complexity) to what they offer.  While there is a substantial amount shared between Rust and C++ that I hope to cover in future posts, one of the 
biggest areas they differ is in the way that they handle “safety” — that is, how the user interacts with memory and mutable data, and how those 
interactions affect program behavior and state. Put simply, Rust aims to be a <em>safe</em> language (with the option to perform tightly-scoped 
unsafe operations via the use of the <code class="language-plaintext highlighter-rouge">unsafe</code> keyword), while C++ is absolutely <em>not</em> a safe language.  Now, to be sure, many in the C++ 
community have recognized the importance of safety — how unsafe code can lead to security vulnerabilities, incorrect results, and unexpected 
program crashes or other behavior.  However, the language itself has incredibly limited support in terms of tools to helping the programmer 
to write safe code. Though there are efforts at laying out best practice <a href="https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines">like the C++ core guidelines</a>,
there is a <strong>huge</strong> qualitative difference between best practice guidelines that a programmer should know and follow, versus safety guarantees
provided by the language and compiler itself.  There are many other differences between these languages, but I’ll stake out the approach to 
safety to be perhaps the most salient.</p>

<h4 id="the-aot-compiled-gc-languages">The AoT, compiled, GC languages</h4>

<p>While they are perhaps less-widely used for tools like the ones I’ve laid out above, there are a host of languages that are still designed 
to provide the computational performance necessary to accomplish such tasks. Here, I’ll group together the most popular ahead-of-time (AoT) 
compiled and garbage collected (GC) languages.  This includes languages like Java (and other JVM languages like Scala and Kotlin), as well as 
different languages that nonetheless adopt the GC approach to memory safety, like Go. There are also languages that mix in a GC with other memory management strategies and provide the user with “opt-in” garbage collection 
(e.g. D and Nim).  However, since people using these languages tend to produce programs that eventually make use of GC somewhat,  I’m grouping them in here,
though technically there are ways to avoid the GC there.</p>

<p>These languages go a very different route that C/C++, and they do, to a large extent, provide important types of memory safety. However, they 
do this at the runtime cost of having a garbage collector, a runtime component of the language that tracks allocated memory and is responsible 
for ensuring that memory is kept alive while it can still be accessed and freed (eventually) when it is no longer in use.  The progress in the 
theory and practice of building scalable garbage collectors has been astounding, and many modern GCs are marvels of engineering. Nonetheless, 
the very presence of a GC imposes a runtime overhead in the presence of heap allocations, and it has been observed that idiomatic reliance on 
the GC for memory management can typically impose a memory overhead of up to 2 times over languages where memory is managed manually (and 
responsibly).  One reason for this (though not the only one) is that such languages tend to make many more heap allocations than languages 
like C/C++/Rust, and store fewer things on the stack. So, the place where these differences show up most frequently when comparing GC’d languages to something like C/C++/Rust is that (a) the 
GC’d languages typically use a constant factor more memory and (b) long-running or memory intensive processes typically incur GC “pauses” 
when the garbage collector kicks in to reclaim large amounts of memory.  Of course, there are several strategies that can be (and often are)
used to mitigate these issues (e.g. like retaining and using small object pools to avoid the garbage collector becoming involved in common and 
reusable allocations), however such strategies often require extra work (sometimes substantial), and are typically not idiomatic in the language.</p>

<p>Therefore, AoT compiled GC’d languages can be quite fast (though typically they take some speed hit compared C/C++/Rust), and they are, in many 
important ways <em>safe</em>, but this often comes at the expense of GC overhead in terms of both runtime and memory usage.</p>

<h4 id="others">Others</h4>

<p>The list above is in no way meant to be comprehensive. Further, as we have been going through a renaissance of types in the development of new 
programming languages, there are also many other (<em>newish</em>) languages that could reasonably by said to occupy a similar space. For example, a language 
like <a href="https://ziglang.org/">Zig</a> aims to be a modern systems-level language, and brings with it some very compelling features.  Since I don’t know too much 
about these other emerging languages (including Zig), I won’t attempt to contrast them too deeply with Rust.  However, I’ll note that while Zig adds features 
on top of C that are a boon for safety, it <em>explicitly</em> does not make the same guarantees or claims on safety as does Rust.  There are also other langauges 
like <a href="https://www.ponylang.io/">Pony</a> of whose existence I am aware, but about which I know even less.  Thus, I won’t be trying to argue in future posts that 
Rust is the <em>only</em> new laguage that is well-suited to developing high-throuhgput bioinformatics tools and methods, but rather that it <strong>is</strong> a great choice 
for this task, and that in the space of such <em>newish</em> languages, it is certainly one of the most widely-adopted and mature.</p>]]></content><author><name>Rob Patro</name></author><category term="blog" /><category term="rust" /><category term="programming" /><category term="bioinformatics" /><category term="tools" /><category term="computer science" /><summary type="html"><![CDATA[ARTE Group at UCAS (LAMP): Autonomous Representation of Temps and Espace, with research in world models, spatial intelligence, and autonomous driving.]]></summary></entry></feed>