Configuring Bazel Module Extensions with tag_class

When writing a Bazel module extension, we sometimes need a way to let the user pass some values to configure things, download binaries from a specified URL, etc. A Bazel module extension itself does not support attributes in the same way repository rules and normal rules do. Instead, it offers another API called tag_class, which lets us define attributes for tags exposed by a module extension.

A single module extension can have multiple tag classes. During extension evaluation, we can iterate over the modules that use our extension and inspect the tags of the class we are interested in.

Downloading a binary

For the purposes of this simplistic article, we will define a very crude repository rule to download any sort of binary. Then we will wire it up through a module extension, which will allow the user to specify the URL and SHA-256 value.

Our repository rule will look like this:

def _binary_download_impl(ctx):
    ctx.download(
        url = ctx.attr.url,
        sha256 = ctx.attr.sha256,
        output = ctx.attr.file_name,
    )

    ctx.file(
        "BUILD.bazel",
        'exports_files(["%s"], visibility = ["//visibility:public"])' % ctx.attr.file_name,
    )

binary_download = repository_rule(
    implementation = _binary_download_impl,
    attrs = {
        "url": attr.string(mandatory = True),
        "sha256": attr.string(mandatory = True),
        "file_name": attr.string(mandatory = True),
    },
)

Then we need to make a module extension with a tag class:

binary = tag_class(
    attrs = {
        "name": attr.string(mandatory = True),
        "url": attr.string(mandatory = True),
        "sha256": attr.string(mandatory = True),
    },
)

def _impl(module_ctx):
    for module in module_ctx.modules:
        for binary_tag in module.tags.binary:
            binary_download(
                name = binary_tag.name,
                url = binary_tag.url,
                sha256 = binary_tag.sha256,
                file_name = binary_tag.name,
            )

binary_download_extension = module_extension(
    implementation = _impl,
    tag_classes = {
        "binary": binary,
    },
)

There are two loops here that are worth pointing out.

module_ctx.modules contains the modules participating in our extension. Each module can then have zero or more binary tags, which we access through module.tags.binary.

For every binary tag we find, we call our repository rule and forward the values from the tag to it.

Finally, the user can configure the extension from their MODULE.bazel:

binary_download = use_extension(
    "//:extensions.bzl",
    "binary_download_extension",
)

binary_download.binary(
    name = "some_binary",
    url = "https://example.com/some_binary",
    sha256 = "...",
)

use_repo(binary_download, "some_binary")

And that’s it.

binary_download.binary(...) creates a tag containing the values we defined in our tag_class. When Bazel evaluates the module extension, _impl receives those tags through module_ctx, and our extension can use them to create repositories.

In this case, the flow is essentially:

MODULE.bazel
    |
    | binary(name, url, sha256)
    v
tag_class
    |
    v
module extension
    |
    | binary_download(...)
    v
repository rule
    |
    v
downloaded binary

Conclusion

This stuff is reasonably well covered in Bazel’s official docs, but I felt the need to try and simplify it further. It is also true that I haven’t come across anything more interesting in the past week. So, looking forward to the next one.

Running Bazel Targets Under Another Executable

When running a binary through Bazel with bazel run, we might find ourselves in a situation where it is desirable to wrap it in a “parent” binary or script. This could be to profile the running binary, fiddle with its environment, run it through a debugger, or any sort of thing really.

To achieve this, we could build the binary and then manually wrap it with whatever we want, but that is a bit tedious and perhaps prone to errors.

Using –run_under

Bazel has a flag specifically for this use case: --run_under. Simply put, it prepends a command to the executable being run.

It is a rather old Bazel feature — --run_under was already around in 2015, during the early days of the public Bazel project.

For example, if we wanted to measure the CPU time of a rules_xcodeproj executable target, we could do something like this:

bazel run //:xcodeproj --run_under="time"

This effectively results in Bazel running something along the lines of:

time <path-to-xcodeproj>

and we get the timing summary once the executable finishes.

The wrapper does not have to be a single command either. --run_under accepts a command prefix, so arguments can be passed to the wrapper as well. For example:

bazel run //:xcodeproj --run_under="time -p"

It is worth noting that --run_under is not specific to bazel run — it can also be used with bazel test.

Going further

Fortunately, this flag is not only able to wrap the binary with something available on the host OS. It can also refer to another executable Bazel target.

This means we can have a *_binary target in our repo and, by referring to it using normal Bazel target label syntax, Bazel will build both targets and use one to wrap the other:

bazel run //:xcodeproj --run_under="//tools:my_wrapper"

Conceptually, this becomes something like:

my_wrapper xcodeproj

This is particularly useful when the wrapper itself is part of the repository and we want Bazel to take care of building it rather than relying on some separately installed host tool.

Conclusion

In my opinion, this is yet another somewhat obscure Bazel feature that can come in handy in certain situations, as it did for me quite recently.

Tree Artifacts Over ZIP Files in Bazel

Typically, when writing Bazel rules, we output files as the final step, and that is what we need most of the time. However, today I want to make a case for outputting directories—or, as Bazel calls them, “tree artifacts.”

There are many reasons to reach for a tree artifact instead of, say, a ZIP file. For one, zipping can be expensive and time-consuming, depending on the size and contents of the output. Every downstream action that needs the directory then has to unzip it again, adding more CPU work, disk I/O, and temporary files.

Archives are also opaque blobs. By contrast, Bazel represents a tree artifact as a directory whose individual files are stored in the content-addressable store. The tree is still treated as a single declared output for dependency and action-cache purposes, but remote caching and remote execution can deduplicate the files it contains instead of repeatedly storing and transferring one large archive.

In the context of Apple-platform apps, we typically need .app bundles during intermediate build steps—not ZIP or .ipa archives. Signing, validation, installation, and other tools naturally operate on the bundle directory. Keeping the .app as a tree artifact means that we spend much less time zipping and unzipping it.

This is especially noticeable when developing locally through rules_xcodeproj, where unnecessary archive and extraction steps slow down the edit-build-run cycle.

There are other advantages too, particularly for remote execution, but it is hard to enumerate them all. The general idea is simple: keep structured outputs as directories for as long as possible, and create an archive only at the boundary where something actually requires one—for example, when producing the final .ipa.

Creating a tree artifact

It is quite easy, and perhaps not worthy of an entire blog post—but here we are.

In the rule implementation function, we simply invoke:

output = ctx.actions.declare_directory(
    ctx.label.name + ".app",
)

The returned value is a File, just like the value returned by ctx.actions.declare_file(), and it can be passed to actions and downstream rules in much the same way:

ctx.actions.run(
    executable = ctx.executable.tool,
    arguments = ["--output", output.path],
    outputs = [output],
)

The action producing the tree artifact must create the declared directory and place all of its contents inside it.

The main thing to remember is that Bazel does not know the tree artifact’s contents during the analysis phase. You cannot inspect the directory and turn arbitrary files inside it into regular declared outputs. Its contents are normally available only at execution time.

One newer option is map_directory, which lets Bazel expand actions over the contents of directories. It is still fairly new, however, and is not yet widely used.

Conclusion

My advice is to always evaluate whether an output really needs to be an archive. If it does not, it is usually better to declare it as a tree artifact.

Exploring map_directory in Bazel 9

In the world of Bazel, we prefer to be explicit as much as possible and register actions ahead of time. In fact, this is not just a preference but a strongly enforced rule that enables all the good stuff Bazel gives us.

However, there is a world where a bit of dynamism is useful—or, in some cases, unavoidable.

Today I am exploring the little-known map_directory API that was introduced in Bazel 9.

Registering actions dynamically

Say we write a rule that generates a couple of files and then want to register a separate Bazel action for each of those files.

The Bazel API did not facilitate this use case before version 9. Sure, we could have resorted to all sorts of tricks and hacks, but there was no nice native way to register actions based on the contents of a generated directory.

This is where map_directory comes in. It allows part of action registration to be deferred until execution time, when the contents of an input directory are known.

map_directory in action

The example below demonstrates generating a couple of files and then copying them with separate actions:

def _copy_files(template_ctx, *, input_directories, output_directories, tools, **_kwargs):
    for src in input_directories["src"].children:
        out = template_ctx.declare_file(src.tree_relative_path, directory = output_directories["out"])
        # Here we register one action per file; Bazel executes the command later.
        # Each action depends only on its input file and the copy tool.
        template_ctx.run(
            executable = tools["copy"],
            inputs = [src],
            outputs = [out],
            arguments = [src.path, out.path],
        )

def _demo_impl(ctx):
    src = ctx.actions.declare_directory(ctx.label.name + "_input")
    out = ctx.actions.declare_directory(ctx.label.name + "_output")

    ctx.actions.run_shell(
        outputs = [src],
        arguments = [src.path],
        command = """
mkdir -p "$1/nested"
echo hello > "$1/a.txt"
echo world > "$1/nested/b.txt"
""",
    )

    copy = ctx.actions.declare_file(ctx.label.name + "_copy.sh")
    ctx.actions.write(copy, '#!/bin/sh\nexec /bin/cp "$@"\n', is_executable = True)

    ctx.actions.map_directory(
        input_directories = {"src": src},
        output_directories = {"out": out},
        tools = {"copy": copy},
        mnemonic = "CopyFile",
        implementation = _copy_files,
    )

    return [DefaultInfo(files = depset([out]))]

map_directory_demo = rule(implementation = _demo_impl)

Creating a target from this rule and building it produces an unremarkable result, but it demonstrates what is now possible with this API at our disposal.

The important part is that _copy_files does not run during the regular analysis phase. It runs later, once Bazel knows the contents of the src tree artifact. At that point, it can inspect its children and register an individual copy action for each file.

Conclusion

While still greatly limited, map_directory allows us to introduce some controlled dynamism into our Bazel builds and makes certain things possible that previously required considerably more creativity.

I am pretty sure there are far more interesting examples of this feature being used, so feel free to look around on Github.

Finally, as always, I suggest consulting the official docs, at least as a reference, as well as the GitHub discussion around dynamic dependencies.

Bazel, APFS Clones, and Disk Space

I was recently writing about the problem of disk space usage when building with Bazel due to its many caches. That led me to pay more attention to the work being done in this area, and I discovered that Bazel 9.3.0 is expected to make better use of copy-on-write cloning on macOS #30776 to avoid unnecessarily duplicating files between the disk cache and the output base.

On APFS, this essentially means copy-on-write: the cloned files initially share the same underlying data blocks, so they don’t immediately take up twice the physical disk space. If either copy is modified, the filesystem only needs to allocate storage for the changed blocks.

Using APFS clones from Swift

Naturally, I got interested in learning how this feature works, and it turns out to be pretty simple. There is a low-level API function, clonefile(...), which does exactly what you would expect: it creates a copy-on-write clone of a file.

If I were to wrap that function in Swift, here is roughly how I would do it. In production code, I would also provide better error handling:

func clone(from source: String, to destination: String) -> Bool {
    let source = source.cString(using: .utf8)
    let destination = destination.cString(using: .utf8)
    let result = clonefile(source, destination, 0)
    return result == 0
}

I recommend taking a look at its man page to learn more about it.

Conclusion

I wanted to share this because I feel like APFS cloning is not talked about enough. It’s a simple API backed by a pretty powerful filesystem feature, and I hope you find it handy someday.

How to Reduce Bazel Disk Space Usage Across Git Worktrees

Ever since the advent of coding agents, a lot of engineers have started utilizing Git worktrees and, by extension, multiple Bazel output bases. We all know the story: Bazel caches tend to take up a lot of disk space when working with large projects. Manual deletion and garbage collection can only take you so far.

Enter bb-clientd

bb-clientd is a daemon that runs on your machine and can act as a local remote cache and proxy for Bazel.

More importantly, it also implements Bazel’s Output Service protocol, introduced in Bazel 7.2. This allows bb_clientd to manage Bazel’s output tree through a virtual filesystem—FUSE on Linux and NFSv4 on macOS—and lazily materialize files when they are actually accessed.

Combined with its content-addressed local cache, this means that multiple Bazel output bases can reuse the same cached content instead of each storing their own copies of identical files. This is particularly useful when working with multiple Git worktrees.

Because bb-clientd already has a nice README, there is no point in me explaining the setup in detail. But here is an example .bazelrc configuration that I keep in my global .bazelrc so I can enable it when needed:

common:bb_clientd --disk_cache=
common:bb_clientd --remote_cache=unix:///Users/<user>/Library/Caches/bb_clientd/grpc
common:bb_clientd --remote_instance_name=local/projects
common:bb_clientd --remote_upload_local_results
common:bb_clientd --experimental_remote_output_service=unix:///Users/<user>/Library/Caches/bb_clientd/grpc
common:bb_clientd --experimental_remote_output_service_output_path_prefix=/Users/<user>/bb_clientd/outputs

This lets me use just the local cache with:

bazel build --config=bb_clientd //...

NOTE: If you’re using rules_swift, make sure to check out the swift.module_map_home_is_cwd feature if you’re not using the Xcode toolchain. It makes module-map generation and compilation assume that header paths are relative to the workspace root, which can be important when using a virtualized output tree.

Conclusion

This is one of those things that I don’t see discussed enough in the Bazel community, yet it can save gigabytes of disk space—especially if you’re regularly working across multiple worktrees.

Preventing Transitive Swift Imports with Bazel

Swift’s permissive nature when it comes to module dependencies has always annoyed me and made my life a bit harder than I expected. Of course, I’m talking about transitive module imports.

You know the situation: there is a ModuleA which is a dependency of ModuleB, and if you declare ModuleB as a dependency of ModuleC, ModuleC can import ModuleA even though it never declared ModuleA as its own direct dependency.

There are a lot of reasons why you might not want that, from making dependencies harder to visualize to enforcing a certain module structure, and so on.

Unfortunately, Swift itself does not provide a mechanism to enforce direct module dependencies—or at least I don’t know of one.

Layering check via rules_swift

One of the reasons I like working with Bazel is the flexibility it provides. One example of that is a feature in rules_swift called layering check, which prevents importing modules that aren’t declared as direct dependencies.

It was introduced a couple of months ago in #1780.

So, for example:

swift_library(
    name = "Bottom",
    srcs = ["Bottom.swift"],
    module_name = "Bottom",
)

swift_library(
    name = "Middle",
    srcs = ["Middle.swift"],
    module_name = "Middle",
    deps = [":Bottom"],
)

swift_library(
    name = "Top",
    srcs = ["Top.swift"],
    module_name = "Top",
    deps = [":Middle"],
)

Now let’s say Top.swift tries to do this:

import Middle
import Bottom

Even though Bottom is available transitively through Middle, Top never declared it as a direct dependency.

With the Swift layering check enabled, building it:

bazel build //:Top --features=swift.layering_check_swift

should fail with an error resembling:

Layering violation in //path/to/package:Top
  The following modules were imported, but they are not direct dependencies:

      Bottom

  Please add the correct 'deps' ...

And that’s exactly what I want.

To fix it, simply declare Bottom as a direct Bazel dependency:

swift_library(
    name = "Top",
    srcs = ["Top.swift"],
    module_name = "Top",
    deps = [
        ":Bottom",
        ":Middle",
    ],
)

Conclusion

This is one of those features I wish more people knew about and used. It gives you much tighter control over your dependency graph and makes it harder for accidental dependencies to get imported.

Profiling Starlark Memory Usage in Bazel

When using Bazel within a large monorepo, there comes a time when memory starts becoming a problem, and it is important to be aware of the available tools that can help with diagnosis. Fortunately, Bazel offers Starlark memory profiling, which we can utilize to see how much memory is being swallowed by the analysis phase.

Setting up

Surprisingly, it is a little tedious to set this up, but in the end, it is not that difficult either.

First, we need to download the allocation instrumenter JAR from Maven Central. Once that’s on disk, we can start Bazel in memory tracking mode like so:

bazel --host_jvm_args=-javaagent:/path/to/java-allocation-instrumenter-3.3.4.jar \
      --host_jvm_args=-DRULE_MEMORY_TRACKER=1 \
      build --nobuild //path/to/target

Here, we start Bazel with memory tracking enabled and pass --nobuild so that only analysis is performed, which is what we want to measure.

After analysis completes, make sure not to shut down the Bazel server because you’ll lose the gathered data.

Now, to produce a memory profile, we’ll use Bazel’s dump command:

bazel --host_jvm_args=-javaagent:/path/to/java-allocation-instrumenter-3.3.4.jar \
      --host_jvm_args=-DRULE_MEMORY_TRACKER=1 \
      dump --skylark_memory=memory_profile.gz

Notice how we repeat the startup flags:

--host_jvm_args=-javaagent:/path/to/java-allocation-instrumenter-3.3.4.jar
--host_jvm_args=-DRULE_MEMORY_TRACKER=1

Failing to do so will result in a Bazel server restart, and you’ll have to repeat the whole process from the beginning.

Dealing with gathered data

Now that you have memory_profile.gz, you need to make sense of it. The Bazel docs recommend using pprof to inspect the profile.

There doesn’t appear to be a pprof formula in the main Homebrew formula index, but installing it is straightforward if you already have Go:

go install github.com/google/pprof@latest

The binary is installed into $GOPATH/bin, which is $HOME/go/bin by default.

Once installed, a useful starting point is a flame graph:

pprof -flame memory_profile.gz

or a text report annotated with source lines:

pprof -text -lines memory_profile.gz

Conclusion

Most Bazel users won’t ever need to gather a memory profile, but I believe it is good to be aware that this exists so you can reach for it when needed. Rule authors should be especially aware of it, as it is not that difficult to make a coding mistake in a rule that results in huge memory usage.

You can find more information about optimizing performance in the official Bazel docs.

Testing Bazel Remote Build Execution Locally with actiond

When configuring remote build execution, it is very important to test changes locally so you don’t waste time on CI. A lot of the time, people don’t even have access to an RBE environment. Yes, BuildBuddy offers their service for free for open-source projects, but I feel like having access to an RBE environment locally is invaluable.

I only recently started utilizing RBE more seriously, and while working through that setup I accidentally discovered actiond. I really wish I had known about it earlier. Having a remote executor that you can run locally makes experimenting with RBE, debugging issues, and checking whether your targets actually execute remotely much easier.

actiond from hermeticbuild

actiond is a full-fledged remote build executor that you can spin up easily simply by downloading it:

curl -L \
  https://github.com/hermeticbuild/actiond/releases/latest/download/darwin-actiond_macos_arm64 \
  -o darwin-actiond_macos_arm64
curl -L \
  https://github.com/hermeticbuild/actiond/releases/latest/download/SHA256.txt \
  -o SHA256.txt
grep ' darwin-actiond_macos_arm64$' SHA256.txt | shasum -a 256 -c -
chmod +x darwin-actiond_macos_arm64

And then just run it:

./darwin-actiond_macos_arm64 serve-vm \
  --listen=127.0.0.1:8980 \
  --root="$HOME/Library/Caches/actiond/vm"

NOTE: I assume you’re running on macOS.

Pointing Bazel at it

It is a matter of setting a few flags, which is best done in .bazelrc, as I can’t imagine anyone manually passing them every time:

build:local_rbe --remote_executor=grpc://127.0.0.1:8980
build:local_rbe --remote_cache=grpc://127.0.0.1:8980
build:local_rbe --platforms=//:linux_x86_64
build:local_rbe --extra_execution_platforms=//:linux_x86_64
build:local_rbe --spawn_strategy=remote
build:local_rbe --genrule_strategy=remote
build:local_rbe --remote_local_fallback=false
build:local_rbe --remote_upload_local_results=false
build:local_rbe --noremote_cache_compression

Then build with bazel build --config=local_rbe //....

Conclusion

This is an incredibly easy way to test whether all your targets build remotely, either when preparing to use a real RBE service or as a rule author wanting to ensure that your rules are fully hermetic and can execute remotely.

I wish I had found actiond earlier, because it removes a lot of the friction from experimenting with RBE locally and makes it much easier to catch remote-execution problems before they reach CI.

Less Boilerplate for Bazel Transitions

I’ve written about Bazel transitions multiple times: first when demonstrating rule extensions, and then when explaining split transitions.

Now I want to showcase a community-built way to apply transitions more easily.

Enter with_cfg.bzl

with_cfg.bzl is a convenient way to apply Bazel transitions. It was created by well-known community member Fabian, and I can’t recommend it enough.

Applying a transition to a plain swift_library

Say we want to ensure that our swift_library is always built with --compilation_mode=opt. We can achieve that quite easily with the following .bzl file:

load("@rules_swift//swift:swift_library.bzl", "swift_library")
load("@with_cfg.bzl", "with_cfg")

opt_swift_library, _opt_swift_library_internal = (
    with_cfg(swift_library)
        .set("compilation_mode", "opt")
        .build()
)

That’s it.

NOTE: You can ignore _opt_swift_library_internal. It needs to be assigned to a global variable because of Bazel’s restrictions on rule definitions, but you aren’t supposed to use it directly. Instead, use the generated opt_swift_library macro.

Finally, load opt_swift_library from the .bzl file containing the code above and use it just like a regular swift_library. The target and all its transitive dependencies will be built with the opt compilation mode.

Prior art

I won’t go into detail about what this would look like without with_cfg.bzl, as that’s already covered in my earlier articles.

Conclusion

Unfortunately, I discovered this way too late, but I’m glad it exists. It makes transitions much less tedious to write and reason about.