• StripedMonkey@lemmy.zip
    link
    fedilink
    arrow-up
    0
    ·
    18 hours ago

    The blame and mistake isn’t honestly on git, but the rest of the tooling ecosystem. This wasn’t a snap decision and the fact that nobody thought to get ready in the 6 god fucking years that this has been planned is the real problem.

    This will be painful, but only because apparently nobody was willing to just start supporting it 5min before the deadline.

  • calcopiritus@lemmy.world
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    I tried to implement my own git from scratch. And when I got to choose a hashing algorithm, I reached the exact same conclusion as the author pretty fast.

    The hashing algorithm doesn’t really matter as long as it’s good enough to prevent accidental collisions. And if there is a collision, it is extremely easy to check.

    If you are about to create a new object and an object with that hash already exists, you compare the 2 objects. If they are the same, nothing happened. If they are different, you just found a collision. Throw some obscure error that will only be witnessed once in the lifetime of the universe and be done with it. Tell the user to change a single bit of the input and be done with it.

    The purpose of object hashes in git was never to provide security. Security is achieved through other means.

    • Kajika@lemmy.ml
      link
      fedilink
      arrow-up
      0
      ·
      edit-2
      17 hours ago

      Security is achieved through other means.

      Let’s explicit this: you can, and should, sign your commits if you want security (no commit tampering). You need to:

      • generated a gpg key: gpg --full-generate-key
      • set git to look for the public key to use: git config --global user.signingkey [public key ID]
      • set git to sign your commits: git config --global commit.gpgsign true

      Note that I used --global here but you can do without to sign on a per-project basis.

      Also gpg UX is terrible, to find the [public key ID] use the command gpg --list-keys, it should looks like:

      pub   ed25519 2026-01-01 [SC] <- type algorithm date and capability (this key is only to Sign and Certify)
            662E3CDD6FE329002D0CA5BB40339DD82B12EF16 <- Public key ID
      uid           [ultimate] my full name (Master Key) <my_email@domain.com> <- owner information (should be your name and email address)
      sub   rsa4096 2026-01-01 [E] <- sub key used to encrypt
      sub   ed25519 2026-01-01 [S] [expires: 2027-01-01] <- subkey used to sign
      

      Last, if you want to save/backup you keys the default location of gpg files is ~/gnupg or you can export keys with gpg --export [key ID] > path/to/file.key for the public part and gpg --export-secret-keys [private key ID] for the private part. Both export can have a -a or --armor argument to output as base64 text instead of raw binary. [private key ID] can be found with gpg --list-secret-keys.

      EDIT: after writing this I checked https://git-scm.com/book/en/v2/Git-Tools-Signing-Your-Work and TIL you can sign tags too.

      • dunj3@feddit.org
        link
        fedilink
        arrow-up
        0
        ·
        12 hours ago

        Worth to keep in mind though that the gpg signature basically only signs the (git) hashes of the contained objects, so the choice of hash function is indeed important for security.

        • Kajika@lemmy.ml
          link
          fedilink
          arrow-up
          0
          ·
          edit-2
          9 hours ago

          Indeed it is still relevant. I am wondering what does it sign for tags.

          EDIT: Are you sure about that? signing the hash felt logical but digging about it this seems a false assumption. git is signing the whole commit object.

          • dunj3@feddit.org
            link
            fedilink
            arrow-up
            0
            ·
            8 hours ago

            Yes, but the “commit object” is just a bunch of metadata that refers to the tree by its hash:

            % git cat-file commit 6de20f6092
            tree 5cb21f9acf8c31ce1a8f6bb0bbc3b9e4e8607ce8
            parent c679abf7d2c0f3f4ee623e0882a50964e9b100cc
            author Junio C Hamano <gitster@pobox.com> 1791306800 -0700
            committer Junio C Hamano <gitster@pobox.com> 1791306812 -0700
            
            The 4th batch
            
            Signed-off-by: Junio C Hamano <gitster@pobox.com>
            

            And the tree in turn refers to the files (or subfolders) by their hashes:

            % git cat-file -p 5cb21f9acf8c31ce1a8f6bb0bbc3b9e4e8607ce8 | head
            100644 blob fd4fb56b6d56789369d4824ad10999369127f5c7	.b4-config
            100644 blob 8168d8a10b3a9e14f6c019e8ffbe5a71dc307d8b	.b4-cover-template
            100644 blob fef04a38402fee6465a6a4225374d493b47421c0	.cirrus.yml
            100644 blob 86b4fe33e5cd98e2a347944559c345419824c245	.clang-format
            100644 blob 82e121a41754b536611c9ce9b2b4aba349e7ed9d	.editorconfig
            100644 blob 26490ad60a74d0968eaf2f77abffcb21d78b18e3	.gitattributes
            040000 tree 5f898fc5ec3429058fd21072db170fff9e49eab4	.github
            100644 blob 0209bd16f209c232748734d911099b37be80fc12	.gitignore
            100644 blob 3f2483550038f6aabb2f8b858b4ee7b29a66072a	.gitlab-ci.yml
            100644 blob cbeebdab7a5e2c6afec338c3534930f569c90f63	.gitmodules
            

            If you can produce a hash collision, you can therefore have two repositories with the same signed commit but different file content.

            • Kajika@lemmy.ml
              link
              fedilink
              arrow-up
              0
              ·
              8 hours ago

              Yes you are right, damned! Why git isn’t signing the diff (patch)? This is so wrong on so many level, signing already hash the content, we can sign multiple GiB without any issue.

    • chris@l.roofo.cc
      link
      fedilink
      arrow-up
      0
      ·
      1 day ago

      Even if it’s not the purpose it can fulfill it. Hashes are used for all kinds of things like traceability or for pinning versions. They are used in a context where people depend on that a hash always has the same code behind it and SHA1 can not guarantee that anymore. A better hashing algo is needed to keep the way how people use git secure.

      It’s not enough to say “you shouldn’t use it that way”. We have to accept the reality of how git is used and adapt.

      • calcopiritus@lemmy.world
        link
        fedilink
        arrow-up
        0
        ·
        1 day ago

        If someone has access to write to a repo you are downloading code from, and you don’t trust that someone, you shouldn’t be running the code in that repo.

        The only “security” vulnerability is: I have audited and verified that the code in commit 05682ab36ca. And I will use only that version.

        What you do then is: download that commit and store it in a local repo, and package it so you can distribute it to your clients.

        Auditing external software is tremendous effort, even if it is open source. Using your local repo instead of downloading that commit from GitHub every time is very little extra effort in comparison.

        And if you really need it. You can always hash with sha256 yourself and verify that the commit with that sha1 still produces the same sh256. So you can keep downloading it each time, just need to store the sha256.

        • chris@l.roofo.cc
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          You are exactly doing what I described. You are saying “your using it wrong” instead of acknowledging that’s how people use it and and then improving that. You might not need this feature because you are “using it correctly” but that doesn’t matter. For the lived reality of many people sha256 is the correct move.

            • chris@l.roofo.cc
              link
              fedilink
              arrow-up
              0
              ·
              23 hours ago

              I have. In some languages you can use git repos as dependencies. It’s good practice to use hashes instead if tags because tags are not immutable. SHA1 hashes are better than tags but better hases would be even better.

              Also traceability. You can build a chain from your commit over ci to the artifact if you do it right. If you can change the files to a hash you can destroy that trust chain. Sha256 makes that impossible.

              People use the hashes in ways you might not. It wasn’t meant as a trust anchor but it became one. And that is now being addressed. I also think that the git maintainers have put a lot of thought into backward compatability and making the change as painless as possible.

              • calcopiritus@lemmy.world
                link
                fedilink
                arrow-up
                0
                ·
                16 hours ago

                Good practice != Something that ensures security. As I said, it’s an even better practice to just host your own fork if what you want is security. The reason to use git hashes is mostly so you don’t have to trust the author following semver in there version number. It’s a matter of ensuring that your code will always compile, not a security feature.

                Have you read the article? Your arguments derive from an assumption of “sha1 is mutable, sha256 is immutable”. First of all, every hashing algorithm is going to have collisions, that’s an unavoidable “feature” of hashes. And sha1 is in no way “mutable”, it takes a great amount of effort (and money) to generate a collision. Furthermore, those collisions are not arbitrary. You have to calculate them beforehand. As the article says: what is more likely? Paying 40k€ in compute time to generate a single collision? Or just paying an open source maintainer 40k€ to let you do 1 new commit that most people are going to download anyway?

    • Scipitie@lemmy.dbzer0.com
      link
      fedilink
      arrow-up
      0
      ·
      1 day ago

      Which are? I’m not joking, “GitHub broken” and “Google has issues” are not arguments against a better encryption default.

      The submodule argument is transitional at best.

      As for the pro side (note that this is my opinion, I’m not deeply involved in the discussion - and with not deeply I mean “not at all”):

      The case the author aggressively ridiculed ("I don’t use GPG to trust, I trust the GitHub authentication!) is one that anyone would have if they’d want to stay safe from centralized systems.

      Another is upstream availability: once an algorithm is identified as broken (no matter how small) usage drops rapidly - there’s a reason why no one uses md5 anymore although it “would be fine” for purely collision mitigation.

      And then there the simple fact that the integrity hash is a security feature, not purely anti collision (as the author also describes). Just because they (nor I!) can’t see an attack vector at the moment doesn’t make this less of an issue.

      In short: it’s the other way around: you’d need way stronger arguments than “there will be a transition pain” to knowingly compromise your security.

      • Wojwo@feddit.online
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        I still like md5 for collision detection, but every damn security scanner flags it. So then you have to add the exception, then explain to auditors about the exception. It’s not worth the hassle for the increased performance.

        • StripedMonkey@lemmy.zip
          link
          fedilink
          arrow-up
          0
          ·
          8 hours ago

          There’s no reason to use md5 if you don’t care about security and only want speed. Non-cryptographic hashes are incredibly fast to calculate and there are far faster hashing algorithms if you want to go that route. Using md5 is a “I don’t want to think hard” answer, and if you don’t want to think hard, then you might as well be secure in your default answer.

  • quantumvoid0@programming.dev
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    thats bad, i got couple of repos on sha-256, i wonder what happens if they drop it, which i think they wont, do i have to like re-create them?