Remote Vivado Builds: More Git, Less Suck

Summary

We've had a "good" build box for a while now, but I struggled to use it for remote builds and regression testing. I tried all the usual ingredients (Parsec, VNC, x2go, sshfs, etc.) a number of times and always found the juice not worth the squeeze. After leaning in for a week or two, I'd just slide back to my old habits (building on a local machine that's barely powerful enough for it.)

This is actually an XY problem. For projects hosted in git, trying to sync or share filesystems with a build server is wrong-headed to start with. That's what git is for. Use it more, not less.

The result is a simple remote-build framework for Vivado projects, using git's built-in "hooks" mechanism. You trigger a build like this:

$ git push -f some-machine:build HEAD:bitstream

or a simulation like this:

$ git push -f some-machine:build HEAD:sim

...or anything else you can write in a bash script. If that's too much typing, use an alias:

$ alias go_bitstream="git push -f some-machine:build HEAD:bitstream"
$ alias go_sim="git push -f some-machine:build HEAD:sim"

and just run go_bitstream or go_sim. The code is at https://github.com/gsmecher/vivado-hooks, alongside a minimal Arty project to try it on.

How It Works

When you push, the remote host kicks off a build using your head-of-tree commit in

some-machine:~/build/autobuild/[commit hash]-bitstream

...using the hooks/bitstream script that's hosted in the source repository. Because it's version-controlled, it's allowed to co-evolve with the project. (That means you can modify hooks without worrying about breaking older versions of your tree.)

The build is hosted in a tmux session (named bitstream-[commit hash]), so it's safe to close your laptop immediately and reconnect to the build later on. (ssh -t build_server tmux attach)

Each build is hosted in its own git worktree, making it addictively easy and safe to launch several builds in parallel, up to the capacity of your build machine to sustain them all. Because build trees are not overwritten or deleted, you maintain a historical record of your builds. (Yes, this gets big, but that's arguably a good use of disk space.)

There are only two ingredients. The first is a post-receive hook installed once on the build server. It doesn't know anything about Vivado; it just maps the branch name you pushed to a script of the same name in the pushed commit's hooks/ directory, and launches it in tmux:

#!/usr/bin/env bash

# This is the "real" post-receive hook that gets installed in
# hooks/post-receive. It is responsible for executing other hooks in this
# directory, and is not useful on its own. See README for details.

BASE_GIT_DIR=$(dirname "$0")/..
unset GIT_DIR

# Check prerequisites
if ! command -v tmux &> /dev/null
then
     echo "ERROR: tmux is not installed. It's an essential prerequisite."
     exit
fi

while read oldrev newrev refname
do
HOOK="$(echo "$refname" | rev | cut -d/ -f1 | rev)"
SESSION="$HOOK-$newrev"
echo Creating tmux session $SESSION
TMP=$(mktemp /tmp/pr-hook.XXXXXX) \
     && git --git-dir="$BASE_GIT_DIR" show "$newrev":"hooks/$HOOK" > "$TMP" \
     && chmod +x "$TMP" \
     && tmux new-session -d -s "$SESSION" "\"$TMP\" \"$oldrev\" \"$newrev\" \"$refname\""
done

The second is the per-target script, which lives in the project. Here is hooks/bitstream: it checks out the pushed commit into a fresh worktree, sources the project's environment, and runs Vivado in batch mode.

#!/usr/bin/env bash

# Check prerequisites
if ! command -v git-lfs &> /dev/null
then
     echo "ERROR: git-lfs is not installed. It's an essential prerequisite."
     exit
fi

oldrev="$1"
newrev="$2"
refname="$3"

WORKTREE=autobuild/$newrev-$(echo "$refname" | rev | cut -d/ -f1 | rev)
echo "Work tree: $WORKTREE"
if ! git worktree add -d "$WORKTREE" "$newrev"
then
     echo "ERROR: git worktree failed."
     exit
fi

cd "$WORKTREE"
echo "Autobuild: working directory is $(pwd)"

source setenv.sh

# Build bitstream
cd tcl
vivado -mode batch -source hw.tcl -source bitstream.tcl

Want a different target? Add another script to hooks/ and push to a branch with the same name.

Why?

Until you've tried this, you probably won't understand how confining your current build process is, or how much it intrinsically conflicts with (or sidesteps) your revision-control processes.

  • Vivado's "remote build" infrastructure relies on a fast shared filesystem, and is clearly intended for corporate LAN environments. This is certainly one type of office, but not the only type of office.
  • Vivado via remote desktop (VNC/x11/etc) is brutal, especially over long distances or narrow pipes, and also requires a shared filesystem or remote copy of your build files.
  • None of these approaches make it convenient to launch a clean, from-scratch build, and they all incentivize you to keep your build tree around long-term (which makes it likely that a clean check-out won't work, and keeps you from casually running multiple builds in parallel.)
  • Running an on-prem or cloud-hosted CI/CD server means one more piece of infrastructure to manage, and pulls at least some build configuration out of your tree.

Since I have begun using this, my workflow has massively re-organized around it.

  • When pushing towards timing closure on a challenging design, it is trivial to run multiple implementation flows in parallel. This allows rapid and principled convergence, rather than endless churn.
  • All design collateral is kept in a clean tree by default, and is trivial to correlate with its commit state (directories are named by hash, after all). It is very simple to produce design surveys and gather cross-linked statistics.
  • "Test builds" and "deployment builds" now all come from the same fast infrastructure.

Regression Testing

The same mechanism runs regression tests, and this is where it earns its keep. Half of effective testbenching is technical (what framework? what should the testbench do?) The other half is "when and how do you use your testbench?"

  • When you find a "live" bug in hardware, do you try to reproduce the bug in your testbench first (and add what was clearly a missing test case)? Or do you leave your testbench out of your workflow (and allow it to wither and die)?
  • Are your testbenches run as part of a regular regression-test process?
  • Do you maintain and run your testbenches as a precondition to merging new code?

A hooks/cosim script makes the answer to the second question "yes" for the price of a git push. In our projects it looks like this:

#!/usr/bin/env bash

unset GIT_DIR

# Check prerequisites
for cmd in git-lfs docker make
do
     if ! command -v "$cmd" &> /dev/null
     then
             echo "ERROR: $cmd is not installed. It's an essential prerequisite."
             exit
     fi
done

oldrev="$1"
newrev="$2"
refname="$3"

WORKTREE=autobuild/$newrev-$(echo "$refname" | rev | cut -d/ -f1 | rev)
echo "Work tree: $WORKTREE"
if ! git worktree add -d "$WORKTREE" "$newrev"
then
     echo "ERROR: git worktree failed."
     exit
fi

cd "$WORKTREE"
echo "Autobuild: working directory is $(pwd)"

source setenv.sh

# Cosimulation targets
set -e  # bail immediately if something fails
cd c/x86_64/cosim
make rtl
make -j8
make test

The make targets hop into a Docker container that captures a Vivado-approved OS, build the RTL into an XSI simulation library, and run the testbenches under pytest using pyxsi. Because tests are discovered and run by pytest, they are parameterized and run in parallel (pytest -n auto), and because xsim is free, there is no seat-counting license nonsense to serialize them through. The runtime of your individual test cases is important (because that affects your quality of life while working on any one test case), but the total runtime of your testbenches is equally important (because you want assurance that no regressions occurred with a relatively short delay.) It seems crazy to me that we accept licensing as a good justification for stacking all of our test cases in series.

Yocto Too

Nothing here is specific to Vivado. We use the same hooks to build Yocto images for our MPSoC targets: crs-yocto carries a hooks/rootfs script that checks out a worktree and runs make rootfs, and the Makefile does the rest (including the Vivado steps that generate the hardware description, and the Docker container they run in). The push looks the same, with tags riding along because the firmware version is derived from git describe:

$ git push -f some-machine:build HEAD:rootfs 'refs/tags/*:refs/tags/*'

A Yocto build takes hours and fills tens of gigabytes, which is exactly the kind of thing you want running in a tmux session on a machine that isn't your laptop.

Getting Results Back

What if timing closure fails? How can I view the results without running into all the same problems (shared filesystems, copied directory trees, ...)?

That's what .dcp files are for. Use the build log to find the most recent checkpoint, and copy it from your build server to your local machine. Vivado lets you open these (open_checkpoint, or via the GUI) without the rest of the project tree.

This is worth a shell function. Because every build tree is named after its commit hash, "the most recent successful build of this collateral" is a one-line ls:

BUILD_SERVER=${BUILD_SERVER:-some-machine}
BUILD_BASE=${BUILD_BASE:-build}

fetch-collateral() {
    # usage: fetch-collateral <collateral> <hash>
    #
    # When hash is not provided, this fetches the most recent successfully
    # built collateral, so it will reach past failed builds. Because the
    # filename has the hash embedded in it, this situation should be fairly
    # noticeable.
    local path hash
    path=$(ssh "$BUILD_SERVER" \
        "cd $BUILD_BASE/autobuild && ls -td ${2}*-bitstream/$1 2>/dev/null | head -n1") || return
    if [ -z "$path" ]; then
        echo "fetch-collateral: no collateral matching $1 for hash ${2:-(unset)} found" >&2
        return 1
    fi
    hash="${path%%-bitstream/*}"
    scp -p "$BUILD_SERVER:$BUILD_BASE/autobuild/$path" "${hash}-$(basename $1)"
}

fetch-dcp() {
    fetch-collateral 'tcl/arty/arty.runs/impl_1/arty_routed.dcp' $1
}

fetch-bit() {
    fetch-collateral 'tcl/arty/arty.runs/impl_1/arty.bit' $1
}

These live in the project's setenv.sh alongside the go_bitstream alias, so they are version-controlled too.

Setting Up

Your build server requires Vivado and some additional packages (tmux, git).

  1. On the build server, set up a bare repository endpoint that will receive your push triggers:

    some-machine$ git clone --bare git@github.com:myorg/myrepo.git build
    
  2. On the build server, install the post-receive hook:

    some-machine$ cd build
    some-machine$ git show main:hooks/post-receive > hooks/post-receive
    some-machine$ chmod 755 hooks/post-receive
    
  3. On your development machine, try it out:

    $ git push -f some-machine:build HEAD:bitstream
    

    ...or just go_bitstream if you set up the alias at the top. (Put this into the repository's setenv.sh script.)

Questions and Answers

Do I need to commit every little change before I push it to the build server?

I don't want millions of "corrected a typo" / "oops" commits in my git tree.

This is why git rebase -i exists. Your local git history can be full of "fixme" commits, and you can force-push these to the build server without any problems. You should, however, rebase (squash, reorder, clarify) this tree before pushing it anywhere else.

I see this as a plus: client-facing builds need to come from properly curated (tagged, rebased, sanitized) git trees, but I also love being able to throw WIP garbage at the build server without anyone else seeing it.

Isn't this just a CI/CD runner with extra steps?

Sort of. A runner (GitHub Actions, GitLab, Jenkins, Buildbot) provides compute and a way to launch containers, and you teach it to call make test and make build at the right moments. A self-hosted runner behind your firewall also solves the license-server problem. If you already have that infrastructure and someone to look after it, use it.

We use GitHub's freebie-tier cloud CI/CD infrastructure for software projects, but a FPGA build server needs a controlled environment and a little more horsepower, hence in-house. I haven't been able to stomach Jenkins (Java) and have bounced off Buildbot a couple of times. With hooks, the "runner" is sshd, the "pipeline" is a bash script in the tree, and the "dashboard" is tmux attach. There is nothing else to install, upgrade, or back up.

Is installing these hooks a security risk?

Well, yes: you are telling your build machine to run a modifiable script. Anyone who can push to the remote machine can run arbitrary code there. However, if you're using an ssh transport, that was probably already true. There is pretty good practical joke potential here.

Do I need a different bare repository on my build server for every project?

Nope. The remote repo can be re-used for any number of projects that do (or don't) have any history in common.

What about tags?

Tagging is a hassle. A plain git push doesn't transfer tags, and --follow-tags only handles annotated ones, so if your build derives a version string from git describe you need to push tags alongside the commit:

$ git push -f some-machine:build HEAD:bitstream 'refs/tags/*:refs/tags/*'

The -f keeps re-pointed tags (e.g. a re-cut release candidate) in sync too.