Git submodule: pull the folders nested inside a repository

A git pull finishes without an error, the superproject moves to the new commit, and the folder holding the submodule is still sitting on last month's code. Or it is empty, with nothing inside it at all. That gap is not a bug and it is not a broken checkout. A submodule is recorded in the superproject as one exact commit id, and pulling the superproject changes that recorded id without touching the files checked out inside the submodule. The work is in deciding which of two different outcomes is actually wanted, and then setting Git up so that outcome happens without anyone having to remember a second command.

What pull changes and what it deliberately leaves alone

The entry for a submodule in the superproject is a gitlink: a path, plus the commit id of a separate repository. It is content, in the same sense that a line of source code is content. When someone else advances the submodule and commits the new id, that commit is what arrives on git pull. The id in the index moves. The submodule's own working tree does not.

Git reports this state clearly once the output is read as intended. Running git submodule status prints the commit id currently checked out in each submodule, the path, and the git describe output for that id. The first character carries the meaning. A - means the submodule has not been initialized. A + means the commit checked out inside the submodule does not match the id recorded in the index of the containing repository. A U means the submodule has merge conflicts. Adding --cached flips the report to show the id the superproject records rather than the one checked out, which is the fastest way to see the two numbers side by side. Adding --recursive walks into nested submodules.

Fetching is a separate axis from checking out, and it is already handled. The fetch.recurseSubmodules setting defaults to on-demand, meaning git fetch and the fetch inside git pull will recurse into a populated submodule when the superproject retrieves a commit that updates that submodule's reference. So after a plain pull, the objects needed are usually already on disk. Nothing has been checked out with them. This is why the fix is almost never a network operation and almost always a local one.

That design is defensible. A submodule can hold uncommitted work, a branch someone is mid way through, or a build artefact. Silently moving its working tree on every pull would throw that away. Git makes the move explicit instead, and the cost of that choice is the confusion this article exists to clear up.

The two questions that look like one

"Update the submodule" means one of two unrelated things, and the commands for them are different.

The first is: put the submodule on the commit the superproject says it should be on. That is reproducibility. Someone committed a specific id, the build was tested against that id, and the local checkout should match. The command is git submodule update, with --init if the submodule has never been set up locally and --recursive if submodules are nested.

The second is: move the submodule to the current tip of its own upstream branch. That is a deliberate change to the superproject, not a sync. The command is git submodule update --remote. The remote used is the branch's remote, from branch.<name>.remote, defaulting to origin. The branch defaults to the remote HEAD unless submodule.<name>.branch is set in .gitmodules or in .git/config, with .git/config taking precedence. To keep the tracking branch state current, update --remote fetches the submodule's remote before working out the target id, which can be suppressed with --no-fetch.

Command Target commit Moves the submodule working tree Submodule HEAD afterwards
git pull in the superproject moves the recorded id in the index no unchanged
git submodule update --init --recursive the id recorded in the superproject yes detached
git submodule update --merge the recorded id, merged into the current branch yes on a branch
git submodule update --remote tip of the submodule's tracked branch yes detached
git submodule update --remote --merge tip of the tracked branch, merged in yes on a branch

The detached HEAD in rows two and four is not a malfunction. The checkout procedure is documented to check out the recorded commit on a detached HEAD, because a commit id is what the superproject stores and a branch name is not. Passing --merge or --rebase keeps HEAD attached, and a merge failure then has to be resolved inside the submodule with the ordinary tools.

Running git submodule update --remote and then committing the result in the superproject is how a submodule bump gets recorded. Skipping the commit means the new id exists only on one machine, and the next person to clone gets the old one.

One config line that removes the second command

Most of the daily friction disappears by setting submodule.recurse. It is a boolean that enables the --recurse-submodules option by default, and it defaults to false.

git config --global submodule.recurse true

The documented reach of that setting is worth knowing precisely, because it is wider than most people assume and narrower than they hope. checkout, fetch, grep, pull, push, read-tree, reset, restore and switch always honour it. clone and ls-files do not. branch honours it only when the experimental submodule.propagateBranches is also enabled. Once it is on, individual commands can opt out with --no-recurse-submodules.

There is one documented trap. Some commands lacking the option will call commands that do honour the setting. git remote update calls git fetch but has no --no-recurse-submodules of its own. The stated workaround is to override the value for a single invocation with git -c submodule.recurse=0.

Setting submodule.recurse also changes two related defaults, which is usually welcome and occasionally surprising. fetch.recurseSubmodules falls back to the value of submodule.recurse when set, turning on-demand into unconditional recursion. And push.recurseSubmodules, which defaults to no, becomes on-demand, so a push that would leave a submodule commit unreachable on the remote gets caught rather than silently succeeding.

Because clone is explicitly not covered, the first clone still needs git clone --recurse-submodules, or a git submodule update --init --recursive immediately after.

When the folder is empty rather than stale

An empty submodule directory means the submodule was never initialized in this clone. git submodule init reads .gitmodules and writes submodule.$name.url into .git/config, using the file as a template. If submodule.active has been configured and no path is given, only the submodules marked active are initialized; otherwise all of them are. git submodule update --init does both steps at once, which is why the explicit init is rarely typed.

URLs change upstream more often than expected, usually when a project moves host or an organisation is renamed. Editing .gitmodules alone is not enough, because init already copied the old URL into .git/config and that is the copy Git uses. git submodule sync --recursive pushes the .gitmodules value back over the configured one. It only affects submodules that already have a URL entry in .git/config, so it is safe to run broadly. To change a URL deliberately, git submodule set-url -- <path> <newurl> writes the new value and synchronises it in one step.

Going the other way, git submodule deinit -- <path> unregisters a submodule: it removes the whole submodule.$name section from .git/config along with the work tree, and later update, foreach and sync runs skip it. Run without a pathspec it errors out rather than deinitialising everything, which is a deliberate guard. Removing a submodule from the repository entirely is a different job, handled by git rm.

Tracking a branch instead of pinning a commit

A superproject can record which branch a submodule is meant to follow, so that update --remote behaves the same way for everyone. The key is submodule.<name>.branch, and the supported way to write it is git submodule set-branch --branch <branch> -- <path>. Passing --default instead removes the key, which makes the tracking branch fall back to the remote HEAD. A special value of . means the submodule should follow the branch with the same name as the current branch in the superproject.

There is a real choice between two mechanisms here, and the documentation states the trade off directly. submodule.<name>.branch travels with the superproject, so every clone agrees on the default upstream branch. branch.<name>.merge inside the submodule gives a more native feel when working in the submodule itself, because a plain git pull there does the expected thing. Running git pull inside the submodule is equivalent to update --remote except for which of those two settings decides the branch name.

Recording a branch does not make the pointer move on its own. The submodule commit id is still content in the superproject, and it still has to be advanced and committed. What the branch setting buys is that everyone's update --remote resolves to the same place.

Seeing the drift before it becomes a merge problem

Submodule changes are easy to commit by accident and easy to miss in review, because the diff is one line containing a commit id. Two settings make them legible.

Setting status.submoduleSummary to true or to a non-zero number makes git status print a summary of commits for modified submodules instead of a bare mention. It defaults to false. Note that the summary is suppressed when diff.ignoreSubmodules is set to all, or for submodules with submodule.<name>.ignore=all, with the exception that status and commit still show staged submodule changes.

Setting diff.submodule changes what a diff shows for a submodule. The default is short, which prints only the commit names at each end of the range. log lists the commits in the range, the way git submodule summary does. diff shows an inline diff of the changed contents. For reviewing a submodule bump, log is the setting that turns an opaque hash pair into a readable change.

For a quick manual audit across every checked out submodule, git submodule foreach evaluates a shell command in each one, with $name, $sm_path, $displaypath, $sha1 and $toplevel available. The example carried in the manual page is exactly the audit most people want:

git submodule foreach 'echo $sm_path `git rev-parse HEAD`'

Submodules defined in the superproject but not checked out are ignored by foreach, so a clean run does not prove every submodule is present. git submodule status is the command that answers that. Keeping that listing beside the folder tree is easier in a file manager with a built-in terminal, where the command and the directory it reports on are visible at once.

Making the update cheap enough to leave switched on

Recursion costs time, and the default settings are conservative. submodule.fetchJobs specifies how many submodules are fetched or cloned at the same time, and it defaults to 1. A positive integer allows that many in parallel, and 0 picks a reasonable default. The same thing can be passed per invocation with git submodule update --jobs <n>. On a repository with several submodules this single change is usually the largest available improvement.

History size is the other cost. git submodule update --depth <n> creates a shallow clone of the submodule. A project can recommend this per submodule through submodule.<name>.shallow in .gitmodules, which the initial clone of a submodule honours by default; --no-recommend-shallow ignores the suggestion. For repositories that are large because of file contents rather than history, git submodule update --filter <filter-spec> applies a partial clone filter to the submodule instead, and clone.filterSubmodules extends a filter given at clone time to submodules. --single-branch limits the update to one branch.

A caution about shallow submodules: a truncated history makes git log and git merge-base inside that submodule behave differently from a full clone, which matters if anyone bisects or builds release notes from it. For a submodule that is only ever compiled, depth 1 is close to free. For one that is actively developed, it is not.

What to change first

Set submodule.recurse to true globally and raise submodule.fetchJobs above 1, then decide per repository whether the pointer should follow the superproject or the submodule's own branch and record that choice with set-branch. That is one afternoon of configuration and it removes the stale-folder problem permanently. Keeping the superproject, its submodules and a shell in one window makes the drift visible while there is still time to act on it, which is the case Atriens is built around.

Frequently asked questions

Why does `git pull` not update my submodules?

Because the submodule's working tree is separate content from the id the superproject records. Pull moves the recorded id in the index, and fetch.recurseSubmodules defaults to on-demand so the objects are usually fetched, but nothing is checked out. Run git submodule update --init --recursive afterwards, or set submodule.recurse to true so pull does it for you.

What is the difference between `git submodule update` and `git submodule update --remote`?

Plain update checks out the commit id the superproject records, which is what reproduces someone else's build. --remote ignores that recorded id and uses the tip of the submodule's remote-tracking branch instead, which is how a submodule gets deliberately bumped. The second one produces a change you then have to commit in the superproject.

Why is my submodule on a detached HEAD after updating?

The default update procedure is documented to check out the recorded commit on a detached HEAD, because the superproject stores a commit id rather than a branch name. Pass --merge or --rebase to keep HEAD attached to the current branch in the submodule. Setting submodule.<name>.update to merge or rebase makes that the default.

The submodule URL changed upstream and my clone still uses the old one. What fixes it?

git submodule init copied the URL from .gitmodules into .git/config, and the copy in .git/config is what Git uses. Run git submodule sync --recursive to push the .gitmodules values back over the configured ones. It only touches submodules that already have a URL entry there, so it is safe to run across the whole repository.

How do I stop submodule updates from being so slow?

Raise submodule.fetchJobs, which defaults to 1, so submodules are fetched in parallel. Then reduce what is transferred: --depth for a shallow submodule clone, submodule.<name>.shallow in .gitmodules to recommend that to everyone, or --filter for a partial clone when the size comes from large files rather than long history.

Does `submodule.recurse` cover every command?

No. checkout, fetch, grep, pull, push, read-tree, reset, restore and switch always honour it. clone and ls-files do not, so the first clone still needs --recurse-submodules. branch honours it only when the experimental submodule.propagateBranches is enabled as well.

Back to all posts