What a confidence score can’t tell you: AI and the authorship of paintings
In 2025, the Badminton House version of Caravaggio’s The Lute Player was submitted to an AI attribution service, which returned a score of 85.7 percent in favor of the attribution. Versions of this scene have become familiar, heralded in headlines announcing that artificial intelligence has settled a question long disputed by art historians. The number looks like science: precise, reproducible, and apparently untouched by the rivalries, liabilities, and habits of judgment that complicate human connoisseurship. Yet the precision of the figure exceeds the clarity of what it represents. What, exactly, has been measured, and how far does that measurement take us toward a historical conclusion about who made the painting? The disagreement surrounding The Lute Player makes the question more than hypothetical.
What follows aims to give art historians, art professionals, and collectors an understanding of these technologies sufficient to evaluate such claims independently. It also situates the tools within the actual work of art history, returning throughout to the first axiom Sonja Drimmer and Christopher Nygren proposed for the discipline’s encounter with AI: “the history of art is not a problem to be solved.” Everything contained in that proposition is important to understanding what computational tools can offer, what kinds of knowledge they can help produce, and where they are categorically unhelpful to the work of attribution.
Authentication, attribution, and inclusion: three terms, one standard of evidence
The question of who made a work of art travels under different names depending on who is asking. Computer-science literature most often uses authentication: a work is evaluated against a target artist, and the system reports whether it belongs to one category or another. Art historians largely speak of attribution: a scholarly judgment about authorship, argued from close looking together with documentary and material evidence, and held open to revision whenever new evidence appears. The catalogue raisonné field speaks of inclusion, although inclusion itself may encompass several degrees of certainty and several ways of describing the relationship between a work and an artist’s established corpus. These vocabularies are not interchangeable. They differ because the intellectual, legal, and institutional stakes surrounding them differ.
Beneath the three vocabularies, however, lies a shared evidentiary expectation, and it is expansive. Any serious judgment of authorship rests on the record of where an object has been and through whose hands it passed, its exhibition history, the account books and correspondence that mention it, and the material analysis that establishes when, where, and from what the object was physically made. Visual examination by expert connoisseurs forms one kind of evidence within that body—indispensable, but one kind among several. The College Art Association’s guidelines make the point explicitly: documentation, stylistic connoisseurship, and technical or scientific analysis are complementary aspects of best practice. Their convergence carries a judgment, in whichever vocabulary it is finally expressed, and supplies the question that should be asked of every AI tool described below: does it add to that body of historical evidence, or does it deliver something else under the name of authentication?
How AI attribution systems work
Many of the AI systems discussed in the news belong to a family of programs called image classifiers. To build one, engineers begin with a collection of digital photographs that experts have sorted into labeled groups. One group may hold works accepted as genuine, while another holds works attributed to imitators, followers, workshop members, or unrelated artists. This labeled collection becomes the training data, and a program is adjusted over many rounds until it can distinguish between the categories represented within it. The adjusted program is the model. When the model receives a new image, it reports the category with which that image is most closely associated, often accompanied by a numerical score expressing the strength of that association. That number is what surfaces in press coverage as a “likelihood,” as in the claim that a painting is 85.7 percent likely to be authentic.
The mechanism measures relationships between an image and the categories defined by the training set. It does not observe an artist at work, reconstruct a chain of ownership, or establish the date of a support. Nor does every AI system used in attribution take the form of an image classifier: some are designed to retrieve documents, compare inscriptions, recognize seals, or organize bodies of visual evidence. The distinction matters because the proper use of a tool depends on understanding what operation it actually performs, rather than accepting the larger historical claim attached to its output.
What AI researchers have actually done
It helps to move past headlines and commercial announcements to the peer-reviewed publications in which researchers have applied computational systems to questions of authorship. The scholars building these systems tend to be considerably more careful about their limits than the coverage that follows them. Three notable cases examine very different bodies of material: the poured paintings of Jackson Pollock, the works of Raphael, and ancient Chinese painting.
A team led by the physicist Richard Taylor at the University of Oregon, with contributions from the Pollock-Krasner Foundation, the Pollock-Krasner Study Center, the International Foundation for Art Research, and Francis O’Connor, assembled digital images of 588 paintings. The researchers divided the images into tiles at multiple scales, allowing the model to compare local patterns across different areas of each composition. Under the conditions of the study, the resulting classifier distinguished the accepted Pollocks from the selected non-Pollocks with a reported accuracy of 98.9 percent. The authors call the output for an individual painting a “Pollock Matching Factor,” a careful name that describes a visual match without, in itself, claiming to establish authorship.
A group led by the computer scientist Hassan Ugail at the University of Bradford, together with the imaging scientist David Stork and colleagues, adapted a model that had already learned general visual features from millions of ordinary photographs. They refined it using 49 accepted Raphael paintings and a comparison group of 49 works by other artists, adding measurements of the fine edges left by the brush. The resulting system achieved 98 percent accuracy on the study’s validation task, a strong result within the particular categories the researchers had defined. When the group applied the system to the Madonna della Rosa, whose degree of workshop participation scholars have long debated, it found most of the composition consistent with Raphael while returning a markedly different result for the face of Joseph. That result did not independently discover the historical conditions of the painting’s production, but it did align with an existing scholarly question about that passage and offered another observation to place beside it.
A third team based at Zhejiang University developed ACPAS, an expert-assistance system for ancient Chinese painting. Rather than issuing a single verdict, the system coordinates image matching, seal recognition, text retrieval, structured databases, and interactive visualization in response to questions posed by a specialist. Its underlying resources include thousands of paintings as well as biographical records, historical texts, and a database of seals, and its design was evaluated with ten experts and researchers. Here the technology is not presented as a replacement for connoisseurship but as an environment in which the connoisseur can assemble and examine evidence more efficiently.
These are serious research projects, and their differences are as instructive as their similarities. The Pollock and Raphael studies classify visual material within defined comparison sets, while ACPAS assists scholars in retrieving and connecting several kinds of evidence. The Raphael researchers state plainly that a full authentication protocol depends on provenance, history, material studies, iconography, condition, and more, and that their system contributes to only one portion of that whole. None of the studies establishes that an image classifier can transform resemblance into historical proof. What they demonstrate is that computational analysis can make particular visual distinctions, organize evidence, and sometimes expose patterns worthy of further scholarly attention.
What the confidence score can and cannot say
The most consequential misunderstanding concerns the confidence score itself. A model’s score, an accuracy figure, and the historical probability of authorship are three different things, although press accounts often allow them to collapse into one another. An accuracy of 98.9 percent describes how frequently a system classified examples correctly across a particular evaluation set, given the labels, comparison works, and testing procedure selected by its researchers. A confidence score concerns the model’s output for one image and expresses the strength with which that image has been assigned to a category. Even in machine learning, such scores are not automatically reliable probabilities: modern neural networks can be highly confident and still poorly calibrated, as research on model calibration has repeatedly shown.
A figure such as 85.7 percent therefore cannot be read, without further demonstration, as an 85.7 percent probability that Caravaggio stood before a particular canvas. At most, it expresses a relationship between the submitted image and patterns the model associates with the category labeled “Caravaggio.” The model has no direct access to the historical event whose probability the number appears to describe. It has access to a digitized image, the training examples, and the labels assigned to those examples. The apparent exactitude of the score risks concealing the distance between those materials and the historical proposition being made.
The question the machine answers is how closely a new image accords with the distinctions it learned from a particular body of examples. That question may be useful, but it differs qualitatively from the question of who made the work. The first concerns measurable features within images; the second concerns a historical event reconstructed through the convergence of visual, documentary, and material evidence. A system trained on a different selection of works, photographed under different conditions, or given a differently composed comparison group may return a different number. Precision within a model should not be mistaken for certainty beyond it.
Training data and the persistence of prior judgment
That last observation points toward the deeper difficulty of the training data. Deep learning ordinarily benefits from enormous quantities of examples, while the accepted corpus of an individual artist is finite, unevenly documented, and often internally diverse. The Pollock study expanded the material available to its model by dividing images into tiles, while the Raphael study relied on fewer than fifty accepted works. Tiling can produce many observations, but it does not turn a limited number of paintings into an equally large number of independent works. The underlying historical corpus remains small, and every choice about which paintings represent the artist continues to matter.
Scarcity of this order creates a familiar technical danger: the model may learn a distinction that works well within the dataset without learning the distinction researchers intended. It may respond to the lighting of a photographic campaign, the tonal effects of a particular reproduction, the dimensions of an image file, or the varnish and condition shared by works photographed in one collection. Machine-learning researchers call this shortcut learning: a model finds an efficient signal that predicts the labels in its training environment but fails when the surrounding conditions change. Such a system can perform impressively in validation and still be art-historically uninformative. The question is therefore not only whether it classifies its test images correctly, but what visual evidence allowed it to do so.
The second problem is epistemological. The works labeled authentic are themselves the accumulated product of prior human attribution, some secure, some contested, and some liable to revision. A model trained on those judgments cannot stand outside them and verify them from a position of independence. It may discover regularities that scholars had not previously articulated, and those regularities may become the beginning of a valuable inquiry, but their meaning remains conditioned by the corpus from which they were learned. One may remove the human from the final moment of classification, yet the human decisions that constituted the training set remain encoded in the categories the model treats as true.
This does not make the model’s observations circular in every respect, nor does it mean that computational analysis can never produce something new. A system may reveal a pattern distributed across many works that no individual scholar had noticed or been able to quantify. The pattern does not, however, interpret itself, and novelty at the level of measurement is not yet novelty at the level of historical knowledge. To become evidence, it must be made legible, tested against the objects and their histories, and incorporated into an argument whose limits other scholars can examine.
Pattern-matching and the long history of forgery
Pattern matching requires particular care because resemblance occupies an uncertain place in the history of attribution. The systems described above can identify regularities of brushwork, texture, edge, color, and pictorial structure that are difficult for the eye to quantify across a large corpus. Some of these regularities may be subtle enough to escape a forger’s attention, and it would be premature to assume that computational comparison can reveal nothing that connoisseurship cannot. Yet the entire history of forgery is also a history of the calculated production of resemblance. A successful forgery is made precisely to reproduce enough of an artist’s visible manner, material character, and historical plausibility to enter the field of accepted works.
The difficulty lies not merely in whether a model is powerful enough to detect patterns, but in whether the patterns learned from its comparison set remain discriminating when confronted with the unknown cases that matter. A classifier tested against works by unrelated artists faces a different problem from one asked to distinguish an autograph work from a sophisticated imitation, a contemporary copy, or the contribution of a highly trained workshop assistant. The “not Raphael” category in the Raphael study, for example, was composed of selected works by other named painters; it could not represent every form that a disputed Raphael attribution might take. Similarly, the Pollock study’s high accuracy describes the works and imitations available to that project, not every forgery that may eventually appear. Greater computational power may refine a visual comparison, but it cannot by itself establish that the comparison group adequately represents the historical problem.
Pattern matching can therefore contribute to attribution without serving as its foundation. A strong match may show that a work deserves closer examination, while a weak match may identify an anomaly that calls for explanation. Neither result establishes what the anomaly means, whether it arises from authorship, chronology, condition, restoration, collaboration, reproduction, or the natural variability of an artist’s practice. The movement from measured difference to historical significance remains an act of scholarship.
What the tools can honestly offer
None of this renders the technology worthless. Responsibly developed with art historians in the room, computational tools can flag a work as anomalous and deserving of scholarly attention, organize and search image corpora at a scale no individual could manage, and surface comparisons a researcher might not otherwise have thought to make. They can help determine where to look, which records to retrieve, and which features require closer examination. These are modest claims beside the promise of automated authentication, but they describe genuinely useful work. They also preserve the proper relationship between the tool and the historical question.
The Chinese painting system points toward one of the more persuasive models: an expert-led process in which the machine helps assemble the materials of historical argument—seals, inscriptions, images, biographical records, and documentary references—while leaving the argument itself to the historian. Developed in response to questions defined by scholars, comparable tools could make provenance records searchable across collections, link documents to the objects they mention, or assemble comparative images that once required months of correspondence and travel to see. Their methods and source collections should remain open to scrutiny, so that their results can be questioned as any other scholarly claim would be. Their uncertainties should be documented rather than compressed into the false clarity of a single authoritative number. At their best, these systems expand access to evidence rather than pronouncing upon it.
What AI cannot responsibly become is an arbiter whose output closes a disputed attribution. Experts may disagree because the documentary record is thin, because the surviving material evidence points in several directions, or because the terms of authorship themselves are complicated by workshop practice and collaboration. A classifier contributes another judgment about visual pattern to that already complex body of evidence. It does not repair the missing link in a provenance, explain a material inconsistency, or decide what kind of participation should count as authorship. Presenting its score as the deciding vote does not resolve the disagreement but conceals the source of it.
The debate surrounding the Badminton House Lute Player offers a useful example. The AI result did not stand outside the existing scholarly dispute and settle it; it entered that dispute as another claim, based on another method, whose relevance and authority still required evaluation. Art historians were not displaced from the question by the appearance of a percentage. They were instead asked to judge what the percentage measured, whether the comparison was historically appropriate, and how much weight the result could bear. That work of judgment is not a residual task left behind after computation has finished but the substance of attribution itself.
The history of art is not a problem to be solved
Authorship is a historical claim that rests on a chain of evidence reaching beyond the object’s photographed surface. Image classifiers measure relationships among visual representations, while retrieval systems can help scholars locate and connect other kinds of evidence. Both may contribute to research, but neither collapses the distance between observation and historical conclusion. At every stage, the distinction must be maintained between what a tool detects and what a scholar can responsibly argue from it. Beneath the individual problems considered here lies a larger question about what the past is and how it becomes known.
The histories we write are arguments built from evidence, woven into and against the work of the scholars who came before us, and revised by those who follow as new evidence and better questions appear. When the work has been done well, the argument holds as evidence accumulates, and the picture grows more exact without ever becoming final. This openness is not a weakness in the discipline’s method, nor is it an uncertainty awaiting technical correction. It is how historical knowledge remains accountable to objects, archives, other scholars, and the conditions under which evidence survives. A confidence score can enter that process as one observation among others, but it cannot bring the process to an end. When it offers to finish what cannot be finished, the offer tells us less about the painting than about how thoroughly the work of art history has been misunderstood.