A scientific model is not a miniature copy of the world. It is a way of answering a particular question by keeping some features in view and setting others aside. A chemist may treat a liquid as perfectly mixed. A physicist may neglect friction. An economist may assume that buyers and sellers share the same information. These choices are not necessarily defects. Often, the model works because it leaves things out.
In many traditional models, at least some of those choices appear in the equations, definitions, or stated assumptions. That can make a weak point easier to inspect: friction was ignored, and friction turned out to matter. Yet older science was never wholly transparent. Observations have always been selected, corrected, cleaned, and fitted; complex simulations can be difficult to understand; and tacit judgments may never reach the page. Philosophers of science therefore treat both idealization and the construction of data as parts of modeling, not as problems invented by artificial intelligence.
The useful contrast is not that equations disclose everything while machine learning discloses nothing. Their simplifying choices often sit in different places and require different forms of inspection.
Here, “artificial intelligence” means data-driven machine learning used in scientific work, especially systems whose behavior is learned from many examples. It does not mean every technology that is marketed as AI.
Consider weather forecasting. Physics-based numerical forecasting and machine-learning forecasting both use measurements such as wind, temperature, pressure, and humidity. A physics-based system calculates how atmospheric conditions evolve under mathematical equations. A system such as GraphCast learns a forecasting relationship from decades of reanalysis data. Neither sees the weather whole. Instruments, geographic coverage, grid resolution, preprocessing, training targets, and scoring rules all help determine what the model can notice.
“Compression” is a useful description of part of this process, but it should not be taken too literally. Machine-learning systems do not all compress data in the same technical sense. More generally, training produces a model that preserves patterns useful for a chosen objective. Information that is rare, absent, or irrelevant to that objective may have little influence, even when it later matters greatly.
The deepest omission may occur before training begins. A dataset is a record of what people and instruments measured; it is not the world itself. A missing population, an unrecorded condition, an unreliable label, or a sensor that fails at the wrong moment cannot be repaired merely by adding more examples of what was already measured. The training data, the model architecture, and the definition of success each narrow reality in a different way.
When a small, explicitly specified model fails, investigators can sometimes point directly to the crack. A system was treated as closed, but an outside force changed it. With a complex learned model, the crack may be distributed across the data, the objective, the architecture, or the conditions in which the model is used. A system can score well on its test data and still fail when the population, instruments, or environment changes. It may also be confident for the wrong reason.
This is a difference of degree, not a law separating two eras. Some traditional simulations are extremely opaque, while a small decision tree or symbolic machine-learning model may be easy to inspect. Age and complexity are not the same thing. The practical questions are how the model was built, which assumptions are documented, and whether its behavior can be tested under the conditions that matter.
Interpretability is one response, but it is not a single method or a guarantee. Researchers may examine which inputs affected an output, probe patterns represented inside a network, compare the system with a simpler model, or design an interpretable model from the start. An explanation must itself be checked: a tidy account produced after a prediction may be persuasive without faithfully describing how the model reached it. The NIST principles for explainable AI accordingly distinguish giving an explanation from giving one that is meaningful, accurate, and aware of its limits.
Explanation is only part of model checking. Researchers can test a model on data excluded from both training and model selection, examine difficult cases, compare performance across conditions, measure uncertainty, vary inputs, and try to reproduce the result. When future conditions may differ from the original data, external, temporal, geographic, subgroup, or distribution-shift tests may also be needed. Peer review can help expose weaknesses in that work, but it cannot substitute for it. A review of interpretable machine learning for scientific discovery identifies data splitting, stability, and uncertainty quantification as central validation problems.
Opacity matters when a model succeeds as well as when it fails. Repeatedly correct predictions can tempt people to treat performance as proof of understanding.
Predictive accuracy is not the same as mechanistic understanding.
A forecast may be valuable even when it does not explain the causes of the event. A model may guide an experiment before anyone can give a clear account of why it works. Those are genuine scientific uses. But if the aim is to identify a mechanism, choose an intervention, or carry a result into unfamiliar conditions, a successful correlation may not be enough.
“Understanding” itself has several meanings. It may refer to identifying a mechanism, showing how changing one cause changes an outcome, connecting a result to a larger theory, or reliably anticipating a new case. A system can contribute to one of these without supplying all of them. The question is therefore not whether AI either understands or fails to understand. It is what kind of knowledge a particular system supports, and what further evidence the scientific claim requires. This concern is developed directly in research on scientific understanding with AI and on AI-created illusions of understanding.
AlphaFold offers a useful example because its achievement and its limits are both concrete. AlphaFold2, developed by Google DeepMind, predicts a protein’s three-dimensional structure from its amino-acid sequence and information derived from related sequences. In the CASP14 assessment, its accuracy was competitive with experimental structures in a majority of cases. More precisely, it addressed the structure-prediction part of the protein-folding problem; it did not provide a complete account of the physical pathway by which every protein folds.
The predictions can guide experiments, suggest biological hypotheses, and make structural information available on a scale that experimental work alone could not quickly reach. They do not make experiments obsolete. Accuracy varies, and an AlphaFold2 prediction may omit ligands, chemical modifications, environmental effects, or the dynamic range of structures a protein can adopt. A 2024 study therefore described these predictions as exceptionally useful hypotheses that can accelerate, but not replace, experimental structure determination.
This is not a compromise verdict. It is an example of science working properly. A learned model produces a result; confidence measures help show where caution is needed; experiments test the structural details; and the disagreements can become new questions. The model is most informative when its output is part of an inquiry rather than the end of one.
Machine learning is not one method, and opacity is not its permanent essence. Some approaches are deliberately interpretable. Others combine learned patterns with equations, conservation laws, or known symmetries. Work on physics-informed machine learning is one attempt to join data with mathematical structure rather than choosing between them.
Such partnerships are often more useful than a contest between “old” and “new.” Machine learning can reveal patterns, accelerate calculations, and suggest hypotheses. Theory can impose constraints and connect a result to mechanisms. Experiment can expose where both are wrong. Which combination is appropriate depends on the question and on the cost of failure.
Artificial intelligence does not abolish scientific simplification. It can relocate some choices into datasets, objectives, and learned internal structure, where they are easier to overlook and harder to trace. Its fluency or accuracy can then create an illusion of completeness. The answer is not to demand that every useful model reproduce the entire world. It is to keep the omissions open to investigation.
The question to ask about any model, old or new, is not only, “Does it work?” It is also: What made it work, what was left out, and what happens when the world departs from the conditions under which it succeeded?
Science: Learning Skillful Medium-Range Global Weather Forecasting
Nature: Highly Accurate Protein Structure Prediction with AlphaFold
Nature Methods: AlphaFold Predictions as Valuable Hypotheses
Annual Review of Statistics and Its Application: Interpretable Machine Learning for Discovery
NIST: Four Principles of Explainable Artificial Intelligence
Nature: Artificial Intelligence and Illusions of Understanding in Scientific Research
This page was created independently of the individuals and organizations discussed, none of whom had editorial control over its contents. It contains no affiliate links, sponsored content, paid placements, or compensated endorsements. Neither the author nor this website received any payment, free or discounted product or service, preferential access, travel, hospitality, gift, or other material benefit connected with this page. Unless expressly disclosed otherwise, neither the author nor this website is affiliated with, sponsored by, endorsed by, or officially connected with any individual or organization mentioned. Names and trademarks are used only to identify the subjects discussed. A mention does not, by itself, constitute a recommendation or endorsement.
Please review the website's full Privacy Policy, Disclaimer, and Further Terms.
For project inquiries, corrections, accessibility assistance, or general questions:
Ardan Michael Blum
345 Forest Avenue
Palo Alto, 94301, California, United States
Telephone: +1 (650) 427-9358
Online: Contact Form