Git Will Stop Working in 2100
My relationship with Git has always had its ups and downs. Since I first created the account, my Git 2026-10-5 15:16:25 Author: hackernoon.com(查看原文) 阅读量:7 收藏

My relationship with Git has always had its ups and downs. Since I first created the account, my GitHub tagline has always been “Doesn’t know how to use git properly”. It is a statement that has maintained its truth, and I expect it to do so for the foreseeable future. Git is a densely packed tool! With this article, you might realize the same tagline applies to most of Git's users too.

My Github bio. "pisanvs - Doesn't know how to use git properly"My Github bio. "pisanvs - Doesn't know how to use git properly"

Git is a fantastic tool. As developers, we often take it for granted. Its basic functions are easy to use, it’s file-based, often doesn’t get in the way, and everyone you work with knows how to use it.

The thing is, the software industry, although pervasive, is narrow. Other industries often rely purely on MS Word’s or Google Docs’ version control… And as someone who has spent most of his short life using Git for keeping track of changes, just looking at the UX for these makes me want to vomit.

Google Docs' version control UIGoogle Docs' version control UI

There are good reasons why developers should share the joy of Git with other industries! This is what I attempted to do recently after years of seeing how my mom took ages comparing almost identical legal documents side by side, tracing laws that amended other laws, and working through lengthy incorporation changes.

I figured it couldn’t be that hard to get my country’s entire legal corpus on Git. After all, lawyers might like the idea of Git and start using it!

Making, breaking and shaping trees

The first challenge I faced was scraping upwards of 330k laws from Chile’s National Congress Library into a Git Repo, programmatically, and before GitHub’s CI timeouts hit.

You see, the Chilean National Congress Library stores all types of normative documents; these include all laws, resolutions, decrees, agreements and others, from 3,780 distinct state bodies. This is what bumps the number of changes up to 330k+, which is, frankly, a lot for Git to handle.

The first challenge was building a graph fitting all of these norms in a simple and ordered representation. My first reflex was to build a DAG, but the reality of law is that everyone either a) works on the same tree or b) works on a completely detached one. So in the end, the DAG ended up looking more like a single line with history being chronologically written from vertex to vertex, no branches, no fun.

I’m not going to dive deep into the details of how the data was scraped and the tree constructed. You can look deeper into it in the repo linked below. The important idea here is that DAGs are used in Git as well. When you write an initial commit in a repository, that’s the beginning of your graph. With the exception of orphan branches, the initial commit is always the ancestor of every single future commit.

In our case, the solution was to use a fast-import stream, which lets you build a git history from a file or stream.

Excerpt from the git-fast-import manpageExcerpt from the git-fast-import manpage

According to the git-fast-import manpage, “fast-import reads a mixed command/data stream from standard input and writes one or more packfiles directly into the current repository. When EOF is received on standard input, fast-import writes out updated branch and tag refs, fully updating the current repository with the newly imported data.”

It’s the perfect solution for a custom tree builder. In our case, we exported the law graph to json, and then built fast-import streams based on “events” that were represented by new norms (in Chile, all modifications are published as new text, with incremental IDs).

This is what part of a fast-import stream looks like, without the actual file data. This is the last commit of the git/git repo.

$ git fast-export --no-data
commit refs/heads/master
mark :2
author Junio C Hamano <[email protected]> 1788366110 -0700
committer Junio C Hamano <[email protected]> 1788366119 -0700
data 66
The 22nd batch

Signed-off-by: Junio C Hamano <[email protected]>
from :1
M 100644 27becd62a941da822a750e62c629f532905f60fa Documentation/RelNotes/2.56.0.adoc

If you tried to export the whole data branch of the ley-chile repo, it would be a 3.1 GB file!

History ends and UNIX begins.

Quickly, Git started showing its wrinkles.

$ git init
Initialized empty Git repository in /home/pisanvs/git-testing/.git/
$ echo "The 50s called." >> historical-test.txt
$ git add historical-test.txt
$ git commit -am "Commit from the past!" --date 1950-01-01
[main (root-commit) f9249d9] Commit from the past!
 Date: Thu Jan 1 14:32:54 2026 -0300
 1 file changed, 1 insertion(+)
 create mode 100644 historical-test.txt


$ GIT_AUTHOR_DATE='1914-06-28 00:00:00 +00:00' GIT_COMMITTER_DATE='1914-06-28 00:00:00 +00:00' git commit -m 'Just in! War started!'
fatal: invalid date format: 1914-06-28 00:00:00 +00:00

Silent failures and misleading errors. Git seems to be in big trouble!

The first attempt shows how committing using the --date option silently failed and reassigned the date to the current year, with the rest of the fields staying the same as now. Is this even a reasonable fallback?

Then, the second attempt, using environment variables to set Git’s date fields to the past, failed with an inaccurate error message: “invalid date format”, when the same date format in 2026 works just fine!

$ touch ww1.txt
$ git add ww1.txt
$ GIT_AUTHOR_DATE='2026-06-28 00:00:00 +00:00' GIT_COMMITTER_DATE='2026-06-28 00:00:00 +00:00' git commit -m 'Just in! War started!'
[main 29e7add] Just in! War started!
 1 file changed, 0 insertions(+), 0 deletions(-)
 create mode 100644 ww1.txt
$ git log
commit 29e7add3556fd9895d7af6c930ec81dd8eb1009e (HEAD -> main)
Author: Max Morel <[email protected]>
Date:   Sun Jun 28 00:00:00 2026 +0000

    Just in! War started!

The “invalid date format” error message can be traced to ident.c, when setting the dates using environment variables.

const char *git_author_info(int flag)
{
    return fmt_ident(getenv("GIT_AUTHOR_NAME"),
                     getenv("GIT_AUTHOR_EMAIL"),
                     WANT_AUTHOR_IDENT,
                     getenv("GIT_AUTHOR_DATE"),
                     flag);
}

const char *fmt_ident(..., const char *date_str, int flag)
{
    ...
    if (date_str && date_str[0]) {
        if (parse_date(date_str, ident) < 0)
            die(_("invalid date format: %s"), date_str);
    }
    ...
}

Where parse_date is a small wrapper function over parse_date_basic that parses datetime formats. Like @<unix> <tz> form, RFC 2822, ISO 8601, YYYY.MM.DD, etc.

Another function called parse_date_force processes human language dates like “2 weeks ago” into a UNIX timestamp, prepended by the @ character.

When either fails to parse the date, on whatever grounds, the message is always “invalid date format”, since die() is never called inside them with an error message, but rather their failures get swallowed by the generic error message one layer above. For the --date case, parse_date_force fails; for the environment variables/ident case parse_date_basic fails.

Avoiding these errors might require the Git inquisitor to step outside of the porcelain and into the plumbing. In my case, I wanted to use fast-import.

Funnily enough, fast-import doesn’t reject negative timestamps! But getting through the back door initially doesn’t mean you won’t eventually run into problems down the line.

Attempting to fast-import an ancient-dated commit makes Git convulse on fsck:

$ git fsck
error in commit 3e670a04...: badDateOverflow: invalid author/committer line - date causes integer overflow
error in commit d2105c64...: badDate: invalid author/committer line - bad date
dangling commit 3e670a04572e74ee93655d03d29153a098cccbdd
exit=1 

Most reasonable hosted Git services run fsck upon receiving pushes. This means that GitHub and others plainly don’t support backdating commits before the UNIX epoch, which is also the reason why my project’s repo has thousands of commits dated 1970-01-01. The actual date is stored on a separate metadata file.

Pushing it

Okay, we know Git can’t handle the past. How about the future? Look, I know, this is the definition of a stretch! But who doesn’t want Git to keep working till the end of time?

Turns out, Git will plainly stop accepting commits after 2099. I’m not kidding.

$ git commit -am "Commit from the future!" --date 2100-01-01
[main 1035302] Commit from the future!
 Date: Thu Jan 1 21:36:18 2026 -0300
 1 file changed, 0 insertions(+), 0 deletions(-)
 create mode 100644 thefuture.lol
$ git commit --amend -am "Commit from the future!" --date 2099-01-01
[main 7d2440d] Commit from the future!
 Date: Thu Jan 1 21:36:35 2099 -0300
 1 file changed, 0 insertions(+), 0 deletions(-)
 create mode 100644 thefuture.lol

This was mind-blowing to me! This didn’t even seem like a reasonable failure mode.

Why is this the case? ecee9d9e. A 21-year-old commit that decided the next century was just too far away to bother calculating leap years.

#define DISABLE_SIGN_COMPARE_WARNINGS

#include "git-compat-util.h"
#include "date.h"
#include "gettext.h"
#include "pager.h"
#include "strbuf.h"

/*
 * This is like mktime, but without normalization of tm_wday and tm_yday.
 */
time_t tm_to_time_t(const struct tm *tm)
{
	static const int mdays[] = {
	    0, 31, 59, 90, 120, 151, 181, 212, 243, 273, 304, 334
	};
	int year = tm->tm_year - 70;
	int month = tm->tm_mon;
	int day = tm->tm_mday;

	if (year < 0 || year > 129) /* algo only works for 1970-2099 */
		return -1;
	if (month < 0 || month > 11) /* array bounds */
		return -1;
	if (month < 2 || (year + 2) % 4)
		day--;
	if (tm->tm_hour < 0 || tm->tm_min < 0 || tm->tm_sec < 0)
		return -1;
	return (year * 365 + (year + 1) / 4 + mdays[month] + day) * 24*60*60UL +
		tm->tm_hour * 60*60 + tm->tm_min * 60 + tm->tm_sec;
}

Both parse_date_basic and parse_date_force call this tm_to_time_t function internally.

But wait! This is just the --date flag and one date format. Will a far-dated commit from environment variables survive a fsck?

$ GIT_AUTHOR_DATE='@68719476736 +0000' git commit -m 'hello, 4147'
$ git log --date=iso
commit 0eca3cdcfb3db1c7c7b526caac5f47dedecd49b9 (HEAD -> main)
Author: Max Morel <[email protected]>
Date:   4147-08-20 07:32:16 +0000

    hello, 4147

...
$ git fsck
# exit = 0

No issues there! Turns out Git survives past 2099, at least when using direct timestamps. All of this begs the question: what else can Git mess up silently?

Git also lies about its identity issues.

Some other findings I made when exploring the option to include author identities for each law were related to the “ident” line on commits.

The first finding was that the error handling in ident.c is pretty much the same as date.c, meaning it’s always somewhere between useless and just completely non-existent. An example of this can be found in the name and email parsing:

GIT_AUTHOR_NAME='a<b>c'  -> stored as [abc]
GIT_AUTHOR_EMAIL='a<b>@x' -> stored as [ab@x]
GIT_AUTHOR_NAME=$'a\nb'  -> stored as [ab]

The problem is that Git never warns about this behavior. So having a name with angle brackets (in my case, when indicating a state organization and the signing author) ends up just stripping them.

This is the same problem class as the --date argument converting the year 1950 to 2026 without notice.

Another finding is in the fast-import stream:

> author n <a@x> <b@x> 100 +0000
fatal: invalid raw date "<b@x> 100 +0000" in ident: n <a@x> <b@x> 100 +0000

Somehow, git blames the author date for a clearly syntax-related issue on the identity line. This is even weirder when looking at the following line that’s accepted with no errors whatsoever:

> author  <> 100 +0000     <- note the double space 

git fsck --strict is clean, so GitHub takes it. git log shows author=[] email=[]. And porcelain refuses to create an identical object: fatal: empty ident name (for <>) not allowed.

On the first issue, the problem is plainly that git strips by default instead of escaping. I wouldn’t think this could be a problem normally, but I recently found out that RFC 5322 allows for emails to contain quoted strings, and escape sequences, like this: “a<hi>@notadomain”@gmail.com, which would get completely stripped and mangled by Git porcelain to ahi@notadomain"@gmail.com which is not a valid email according to the RFCs.

When using fast-import, we get the following error:

fatal: missing > in ident string: Max <"a<hi>@notadomain"@gmail.com> 100 +0000

Again, a misleading error. Git's ident grammar is:

name SP "<" email ">" SP timestamp SP tz

< and > are the only delimiters, and the field between them is positional. RFC 5322 explicitly permits <, >, @ and spaces inside a quoted local part… This means that Git’s email field is not necessarily an email address. It doesn’t even check for an @ character’s existence.

Git, weirdly, has a lot of duplicated and reimplemented functionality, which wouldn’t be a big problem if it wasn’t for the fact that no two parsers agree on what to do! I’m sure there are reasons why having a single parser for everything might not be what the maintainers want. But seeing so much of this non-deterministic behavior is worrying, at least to me, in my very weird use cases.

Git is not made for this

After playing around with Git to this point, it has become obvious that git is not made for this purpose at all. This was my first time poking around git internals, and looking through the code. The jankiness, the trust-me-bros, and the 20-year-old pending TODOs were intriguing. I’m starting to question my level of trust in Git, yet I’m too emotionally bound to it to just break up… It’s effectively Git saying the age-old “it’s not you, it’s me.”

I’ll continue to use git for the foreseeable future. It’s feature-complete for all of my work. This one exception is a use case in which, looking back, I should never have used git. Still, version control is something that coders have been refining since the 70’s. It’s weird to still have these quirks.

Finding solutions and alternatives

My solution for these issues in the meantime was to store metadata files alongside the norms and build a web UI around them that implements most of git’s usefulness in a friendly interface. I still kept the promise of having chronologically ordered and dated commits, just that before 1970, all commits had to be dated one second apart from each other, to at least enable sorting them.

I tested Git, Mercurial, and Fossil with historical commit dates. Git rejected pre-1970 commits through its normal commit path as we’ve seen. Mercurial, instead, accepted dates all the way back to the signed 32-bit-floor. This means it accepts commit timestamps only from 1901-12-13 20:45:52 UTC through 2038-01-19 03:14:07 UTC, the inclusive signed-32-bit Unix-time range; dates one second outside it are rejected rather than clamped or wrapped. This is actually a bit better than Git, spanning about 136 years, instead of 129. But, a bit less future-facing, given we won’t be able to commit in mercurial after 2038. (in less than 12 years!)

It is important to note, though, Mercurial only has this issue because of a non-conditional guard that gets triggered on dates with values larger than the 32-bit span. This was justified in a 2005 commit, arguing for compatibility with 32-bit runtimes. To keep a 21-year-old compatibility check in a world where most of everyone has already deprecated 32-bit computing.

The only option that stood the literal test of time was Fossil. It allowed storage of pretty much any date value, accepting the first commit of the corpus dated 1855. The actual limit of the julianday float is November 24, 4714 B.C. at noon UTC. Fossil rendered check-ins on all dates. This is because Fossil works with a SQLite backend and uses a REAL column with julianday values stored inside it.

For raw historical dates, the very aptly named Fossil takes the crown. It is the only one of the three that could handle the full date range for our corpus cleanly.

Conclusion: In Linus’ words

This excerpt is from the git README. It is very accurate; I don't have much else to add:

The name "git" was given by Linus Torvalds when he wrote the very first version. He described the tool as "the stupid content tracker" and the name as (depending on your mood):

  • random three-letter combination that is pronounceable, and not actually used by any common UNIX command. The fact that it is a mispronunciation of "get" may or may not be relevant.
  • stupid. contemptible and despicable. simple. Take your pick from the dictionary of slang.
  • "global information tracker": you're in a good mood, and it actually works for you. Angels sing, and a light suddenly fills the room.
  • "goddamn idiotic truckload of sh*t": when it breaks

References

This article was published under HackerNoon’s Business Blogging program


文章来源: https://hackernoon.com/git-will-stop-working-in-2100?source=rss
如有侵权请联系:admin#unsafe.sh