{"article":{"slug":"science-is-open-software","title":"Science is open software","subtitle":null,"summary":"Jens Egholm Pedersen argues modern science is synonymous with open-source software: why reproducible code matters as much as papers, and what researchers should do next.","content_type":"essay","language":"en","canonical_url":"https://jepedersen.dk/blog/202505_research/","author":{"name":"Jens Egholm Pedersen","url":"https://jepedersen.dk/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Jens Egholm Pedersen","url":"https://jepedersen.dk/","listing_slug":null,"listing":null},"topics":[{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1517,"reading_minutes":7,"published_at":"2025-05-24T00:00:00.000Z","added_at":"2026-09-19T03:08:49.474Z","updated_at":"2026-09-19T03:08:49.474Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/science-is-open-software","markdown_url":"https://listedarticles.com/articles/science-is-open-software.md","example":false,"citation":"Jens Egholm Pedersen, Jens Egholm Pedersen. \"Science is open software.\" 24 May 2025. https://jepedersen.dk/blog/202505_research/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://jepedersen.dk/blog/202505_research/"},"body_markdown":"# Science is open software\n\n**TL;DR** I claim that modern science is synonymous with open source software.\nThis post explains why, why it matters, and what you can (and should) do next.\n\nWhy do you care about (open source) software? - Everyone\n\nI spend a *lot* of my time working on software.\nI have been asked why software matters more times than I can remember.\nSoftware is, people say, not science.\nIt’s a time sink, something to rush past in the pursuit of what really matters: results (and papers if you’re in academia).\nPublish or perish.\n\nWell. I think software matters. In fact, I think open source software *is* science. Or, at least computational science.\nAnd this post tells you why.\nWhy we as scientists must insist on the scientific method and why that means working on open and reproducible software.\nThis post is not easy to write.\nIt challenges many of the current trends in academia, but it is an important move *towards better science* that doesn’t turn us all *insane*.\n\n## What is science?\n\nIf you look up science\n\n on Wikipedia, here’s what hits you:\n\nScience is a systematic discipline that builds and organises knowledge in the form of testable hypotheses and predictions about the universe. - Wikipedia\n\nNow, go and grab a random arXiv paper.\nIt clearly contains “knowledge” of some sort.\nBut, does the paper contribute __predictions__ that are __testable__ and can by __systematically organized__?\n**Can you test it?**\n**Can you systematize it?**\n\nThe answer is never a flat no\n\n, but it’s hard.\nYou rarely have direct access to that knowledge.\n\n### The good explanation - inner models\n\nIf the organism carries a 'small-scale model' of external reality and of its own possible actions within its head, it is able to try out various alternatives, ... and in every way to react in a much fuller, safer, and more competent manner to the emergencies which face it.\n\nIn his *excellent* book The Nature of Explanation\n\n, Kenneth James Williams Craik posits that we use small simulations\n\n of reality to explain and predict the world outside.\n\nThis point seems obvious today, but it highlights the goal of pursuing science in the first place:\nyou, as an acting entity, improves your **inner model** to the point that you can make *better* predictions than before.\nThe **inner model** here is critical: if the arXiv paper does not help their readers predict the world, it is not science.\nThis is why computational reproducibility matters–software is how we encode and share predictive models.\n\n## What is reproducibility?\n\nRecall that according to Wikipedia, it is *not* enough to demonstrate results alone.\nResults have to be (1) systematic and they have to be (2) testable.\n\nIt is entirely possible that the given paper is too hard to understand or unaccessible to the audience for other reasons.\nThat does not mean that there are no scientific insights to find—readers may find ways to systematize them on their second or third reading.\nNo, it means that *you specifically* cannot take the idea as your own, test it, and use it to improve your world model.\n\nReproducibility, in this context, is not only the duplication of results. It is the ability to take the scientific idea, embed it into your own inner model, adapt it, and build upon it—or discard it because it reduces predictabilitly.\n\nIf an idea is not reproducible, the findings cannot be expanded. And are, therefore, useless.\n\nThis becomes clear if we do a quick thought-experiment where we replace “software model” with “mathematical model”.\nJust as we wouldn’t accept a physics paper that said our equations predict X but we won’t show the math\n\n, we shouldn’t accept (computational) science that hides its methods.\n\n## Why is software science?\n\nHow many fields have been held back, and how many people have had their careers disrupted, because of a buggy program? - Greg Wilson\n\n\nSoftware is ubiquitous in modern science. Anything from CoVid models to search algorithms to lab protocols are build on software built by other people. Researchers are busy people. They don’t bother to look through all software dependencies to verify correctness, understand implementation details, or check for potential errors that could invalidate results.\n\nFrom that follows that the **scientific results depend on the software**.\nIf the software is wrong, the science is wrong.\n(Software bugs already cause numerous retractions, such as here, here, here, and several places here).\n\nAnd that is well and good, because at some point we *have* to trust and rely on other’s work.\nFor that to happen, it (software) needs to be reliable.\n\n## Why open source?\n\nWe found that software needs to be\n\n1. **Reproducible** , meaning executable, as well as modifiable, and\n2. **Reliable** , meaning that the results are consistently trustworthy\n\nModifiability is important for science for the same reason that equations are important for scientific predictions.\nReliability is crucial because we want *systematic* improvement of our knowledge, not flaky and partial results that only work occasionally.\n\nThis is what open source software gives us.\nWe can change code and retrofit it to suit our needs (just think about Hugging Face models) and we can iterate upon it to continue to improve it.\nIt already *generates trillions in value* and there is room for much, much more.\n\nOf course, open source software is not a perfect cure.\nThere are IP and security concerns, bugs can still occur, and stability can be a problem.\nBut at least the imperfections are on *public record*.\nThey can be amended and improved, just like our scientific understanding.\nFrom that perspective, one can claim that open source software *is* the scientific method—just in simulation.\n\n## A vision for future science\n\nIf we accept these premises we can ask: what would truly open (computational) science look like?\n\n**Every result is instantly reproducible.**\nWhen you read a paper claiming that a new drug reduces symptoms by 30%, you click a link and watch the exact analysis run in your browser. The data processing, statistical tests, and visualizations execute in seconds using the same environment the authors used—preserved perfectly through reproducible containers.\n\n**Scientific software evolves like Wikipedia.**\nClimate models aren’t developed in isolation by single labs, but maintained by global communities. When a researcher in Kenya discovers a bug in atmospheric turbulence calculations, the fix propagates instantly to climate simulations worldwide. Models improve continuously rather than languishing in academic silos.\n\n**The pace of discovery accelerates.** Instead of each researcher building from scratch, we stand on shoulders of giants whose work is not just readable, but runnable and modifiable. Scientific progress compounds at an unprecedented rate.\n\n**Trust in science strengthens.** When climate models, economic forecasts, and medical recommendations are built on transparent, auditable code, public confidence grows. Science communication improves because the models themselves become part of the conversation—not just their conclusions.\n\nThis isn’t utopian fantasy. Every piece already exists—open source communities, reproducible environments, collaborative development platforms. We just need to shape them into a coherent vision for how science should work in the digital age.\n\nThe question isn’t whether this future is possible. The question is: how quickly can we build it?\n\n## What now?\n\nI posit that open source software is a necessary condition if we are to science in a computerized world.\nSoftware is executable mathematical models\n\n that we should prioritize much higher.\n\nWe still have work to do and this is how you can help:\n\n- **Share and document your code**\n  - Papers without code is less scientific because it is harder to build on the insights. In the ideal world any claim should be backed up by *reproducible* code. Always use code from day 1 and always share it.\n- Papers without code is less scientific because it is harder to build on the insights. In the ideal world any claim should be backed up by \n- **Write stable code, use NixOS**\n  - Code should be reliable and work in perpetuity. That means making sure dependencies and environments are kept constant. The best way to do that is to use *reproducible environments* . NixOS is quickly becomming the biggest and best tool there is. It will guarantee that your code will run*exactly the same way* , even 100 years in the future. Docker, Conda, and similar tools are better, but NixOS gives more comprehensive guarantees.\n- Code should be reliable and work in perpetuity. That means making sure dependencies and environments are kept constant. The best way to do that is to use \n- **Build on existing tools instead of creating your own**\n  - For the common scientific knowledgebase to improve, we need cross-platform tools. This is particularly true for small fields such as neuromorphics, where a recent Nature paper pointed out that open source software is key to scaling. Go check out the Open Neuromorphic software guide and see if you can’t find libraries close to your work.\n- **Promote academics that work on software**\n  - Given the huge importance of code, Academic promotions should value software contributions\n\nThe scientific revolution succeeded because it insisted on transparency, reproducibility, and constant scrutiny. The open source movement embodies these same principles for software, but there is much more work to be done.\n\nWill you help make software scientific?","body_html":"<h1 id=\"science-is-open-software\">Science is open software</h1>\n<p><strong>TL;DR</strong> I claim that modern science is synonymous with open source software.\nThis post explains why, why it matters, and what you can (and should) do next.</p>\n<p>Why do you care about (open source) software? - Everyone</p>\n<p>I spend a <em>lot</em> of my time working on software.\nI have been asked why software matters more times than I can remember.\nSoftware is, people say, not science.\nIt’s a time sink, something to rush past in the pursuit of what really matters: results (and papers if you’re in academia).\nPublish or perish.</p>\n<p>Well. I think software matters. In fact, I think open source software <em>is</em> science. Or, at least computational science.\nAnd this post tells you why.\nWhy we as scientists must insist on the scientific method and why that means working on open and reproducible software.\nThis post is not easy to write.\nIt challenges many of the current trends in academia, but it is an important move <em>towards better science</em> that doesn’t turn us all <em>insane</em>.</p>\n<h2 id=\"what-is-science\">What is science?</h2>\n<p>If you look up science</p>\n<p> on Wikipedia, here’s what hits you:</p>\n<p>Science is a systematic discipline that builds and organises knowledge in the form of testable hypotheses and predictions about the universe. - Wikipedia</p>\n<p>Now, go and grab a random arXiv paper.\nIt clearly contains “knowledge” of some sort.\nBut, does the paper contribute <strong>predictions</strong> that are <strong>testable</strong> and can by <strong>systematically organized</strong>?\n<strong>Can you test it?</strong>\n<strong>Can you systematize it?</strong></p>\n<p>The answer is never a flat no</p>\n<p>, but it’s hard.\nYou rarely have direct access to that knowledge.</p>\n<h3 id=\"the-good-explanation-inner-models\">The good explanation - inner models</h3>\n<p>If the organism carries a &#39;small-scale model&#39; of external reality and of its own possible actions within its head, it is able to try out various alternatives, ... and in every way to react in a much fuller, safer, and more competent manner to the emergencies which face it.</p>\n<p>In his <em>excellent</em> book The Nature of Explanation</p>\n<p>, Kenneth James Williams Craik posits that we use small simulations</p>\n<p> of reality to explain and predict the world outside.</p>\n<p>This point seems obvious today, but it highlights the goal of pursuing science in the first place:\nyou, as an acting entity, improves your <strong>inner model</strong> to the point that you can make <em>better</em> predictions than before.\nThe <strong>inner model</strong> here is critical: if the arXiv paper does not help their readers predict the world, it is not science.\nThis is why computational reproducibility matters–software is how we encode and share predictive models.</p>\n<h2 id=\"what-is-reproducibility\">What is reproducibility?</h2>\n<p>Recall that according to Wikipedia, it is <em>not</em> enough to demonstrate results alone.\nResults have to be (1) systematic and they have to be (2) testable.</p>\n<p>It is entirely possible that the given paper is too hard to understand or unaccessible to the audience for other reasons.\nThat does not mean that there are no scientific insights to find—readers may find ways to systematize them on their second or third reading.\nNo, it means that <em>you specifically</em> cannot take the idea as your own, test it, and use it to improve your world model.</p>\n<p>Reproducibility, in this context, is not only the duplication of results. It is the ability to take the scientific idea, embed it into your own inner model, adapt it, and build upon it—or discard it because it reduces predictabilitly.</p>\n<p>If an idea is not reproducible, the findings cannot be expanded. And are, therefore, useless.</p>\n<p>This becomes clear if we do a quick thought-experiment where we replace “software model” with “mathematical model”.\nJust as we wouldn’t accept a physics paper that said our equations predict X but we won’t show the math</p>\n<p>, we shouldn’t accept (computational) science that hides its methods.</p>\n<h2 id=\"why-is-software-science\">Why is software science?</h2>\n<p>How many fields have been held back, and how many people have had their careers disrupted, because of a buggy program? - Greg Wilson</p>\n<p>Software is ubiquitous in modern science. Anything from CoVid models to search algorithms to lab protocols are build on software built by other people. Researchers are busy people. They don’t bother to look through all software dependencies to verify correctness, understand implementation details, or check for potential errors that could invalidate results.</p>\n<p>From that follows that the <strong>scientific results depend on the software</strong>.\nIf the software is wrong, the science is wrong.\n(Software bugs already cause numerous retractions, such as here, here, here, and several places here).</p>\n<p>And that is well and good, because at some point we <em>have</em> to trust and rely on other’s work.\nFor that to happen, it (software) needs to be reliable.</p>\n<h2 id=\"why-open-source\">Why open source?</h2>\n<p>We found that software needs to be</p>\n<ol><li><strong>Reproducible</strong> , meaning executable, as well as modifiable, and</li><li><strong>Reliable</strong> , meaning that the results are consistently trustworthy</li></ol>\n<p>Modifiability is important for science for the same reason that equations are important for scientific predictions.\nReliability is crucial because we want <em>systematic</em> improvement of our knowledge, not flaky and partial results that only work occasionally.</p>\n<p>This is what open source software gives us.\nWe can change code and retrofit it to suit our needs (just think about Hugging Face models) and we can iterate upon it to continue to improve it.\nIt already <em>generates trillions in value</em> and there is room for much, much more.</p>\n<p>Of course, open source software is not a perfect cure.\nThere are IP and security concerns, bugs can still occur, and stability can be a problem.\nBut at least the imperfections are on <em>public record</em>.\nThey can be amended and improved, just like our scientific understanding.\nFrom that perspective, one can claim that open source software <em>is</em> the scientific method—just in simulation.</p>\n<h2 id=\"a-vision-for-future-science\">A vision for future science</h2>\n<p>If we accept these premises we can ask: what would truly open (computational) science look like?</p>\n<p><strong>Every result is instantly reproducible.</strong>\nWhen you read a paper claiming that a new drug reduces symptoms by 30%, you click a link and watch the exact analysis run in your browser. The data processing, statistical tests, and visualizations execute in seconds using the same environment the authors used—preserved perfectly through reproducible containers.</p>\n<p><strong>Scientific software evolves like Wikipedia.</strong>\nClimate models aren’t developed in isolation by single labs, but maintained by global communities. When a researcher in Kenya discovers a bug in atmospheric turbulence calculations, the fix propagates instantly to climate simulations worldwide. Models improve continuously rather than languishing in academic silos.</p>\n<p><strong>The pace of discovery accelerates.</strong> Instead of each researcher building from scratch, we stand on shoulders of giants whose work is not just readable, but runnable and modifiable. Scientific progress compounds at an unprecedented rate.</p>\n<p><strong>Trust in science strengthens.</strong> When climate models, economic forecasts, and medical recommendations are built on transparent, auditable code, public confidence grows. Science communication improves because the models themselves become part of the conversation—not just their conclusions.</p>\n<p>This isn’t utopian fantasy. Every piece already exists—open source communities, reproducible environments, collaborative development platforms. We just need to shape them into a coherent vision for how science should work in the digital age.</p>\n<p>The question isn’t whether this future is possible. The question is: how quickly can we build it?</p>\n<h2 id=\"what-now\">What now?</h2>\n<p>I posit that open source software is a necessary condition if we are to science in a computerized world.\nSoftware is executable mathematical models</p>\n<p> that we should prioritize much higher.</p>\n<p>We still have work to do and this is how you can help:</p>\n<ul><li><strong>Share and document your code</strong><ul><li>Papers without code is less scientific because it is harder to build on the insights. In the ideal world any claim should be backed up by <em>reproducible</em> code. Always use code from day 1 and always share it.</li></ul></li><li>Papers without code is less scientific because it is harder to build on the insights. In the ideal world any claim should be backed up by </li><li><strong>Write stable code, use NixOS</strong><ul><li>Code should be reliable and work in perpetuity. That means making sure dependencies and environments are kept constant. The best way to do that is to use <em>reproducible environments</em> . NixOS is quickly becomming the biggest and best tool there is. It will guarantee that your code will run<em>exactly the same way</em> , even 100 years in the future. Docker, Conda, and similar tools are better, but NixOS gives more comprehensive guarantees.</li></ul></li><li>Code should be reliable and work in perpetuity. That means making sure dependencies and environments are kept constant. The best way to do that is to use </li><li><strong>Build on existing tools instead of creating your own</strong><ul><li>For the common scientific knowledgebase to improve, we need cross-platform tools. This is particularly true for small fields such as neuromorphics, where a recent Nature paper pointed out that open source software is key to scaling. Go check out the Open Neuromorphic software guide and see if you can’t find libraries close to your work.</li></ul></li><li><strong>Promote academics that work on software</strong><ul><li>Given the huge importance of code, Academic promotions should value software contributions</li></ul></li></ul>\n<p>The scientific revolution succeeded because it insisted on transparency, reproducibility, and constant scrutiny. The open source movement embodies these same principles for software, but there is much more work to be done.</p>\n<p>Will you help make software scientific?</p>","headings":[{"level":1,"text":"Science is open software","id":"science-is-open-software"},{"level":2,"text":"What is science?","id":"what-is-science"},{"level":3,"text":"The good explanation - inner models","id":"the-good-explanation-inner-models"},{"level":2,"text":"What is reproducibility?","id":"what-is-reproducibility"},{"level":2,"text":"Why is software science?","id":"why-is-software-science"},{"level":2,"text":"Why open source?","id":"why-open-source"},{"level":2,"text":"A vision for future science","id":"a-vision-for-future-science"},{"level":2,"text":"What now?","id":"what-now"}]}}