Skip to content

Artificial General Intelligence for Law

Angus McLeod

An Artificial General Intelligence for Law (AGIL) is a jurisprudentially-grounded neuro-symbolic artificial intelligence (AI) system that produces auditable and accurate legal reasoning. An AGIL has a core of symbolic logic which reasons over formal representations of norms and facts. Neural networks, particularly Large Language Models (LLMs), add scale and flexibility and in return the symbolic core scopes and verifies their inputs and outputs. This symbiotic balance of symbolic accuracy and neural breadth allows for auditable non-monotonic legal reasoning at scale. An AGIL is positivist insofar as what counts as law depends on social facts as normalised across jurisdictions by a weakly universal system design. It is argued that an AGIL of this nature will help to address the long-standing problems created by the complexity and volume of laws in modern legal systems. It is further argued that the project of legal alignment of AI (‘legal alignment’) needs both a jurisprudentially grounded concept of law and an empirically grounded technical project if it is to work. It is argued that legal alignment should be understood, like the safety of autonomous driving systems, as an emergent property of an evolving sociotechnical system.

Contents

1. Law has a problem

Year-on-year, decade-on-decade, the world and life seems to become more complex (with new technologies, new communities, new demands, new individual and societal frictions, and all at a faster and faster pace). Meanwhile, the law seems to have grown like ‘Topsy’: the algorithms and manifestations of the law have multiplied exponentially and become ever more complex and voluminous. The fact is, we labour under the heavy yoke of a lot of law and a lot of dense, complex law at that.1

1.1 Law is overly complex

Over the course of the last hundred years law has grown increasingly complex. An illustrative case study is the European Union’s (EU) General Data Protection Regulation (GDPR).2 Even a brief survey of the GDPR’s vital statistics has a soporific power. It runs to 99 articles and 173 recitals, occupying around 88 pages in each of the EU’s 24 official languages. The GDPR’s 50-odd opening clauses leave multiple contentious questions to national specification, and each of the EU’s 27 member states has enacted its own implementing legislation, with Germany’s BundesdatenschutzgesetzFederal Data Protection Act running to 86 sections. The European Data Protection Board (EDPB) has issued more than 60 sets of guidelines, more than 250 opinions under the GDPR’s consistency mechanism, and dozens more recommendations, statements, and joint opinions with the European Data Protection Supervisor (a different body from the EDPB), while the working party documents that pre-dated the EDPB are still potentially relevant.3 The supervisory authorities of the 27 member states add their own corpora of guidance, codes of practice, sanctions decisions, and informal communications. Germany alone has 18 federal and state data protection authorities, which theoretically coordinate their positions through the DatenschutzkonferenzData Protection Conference . National authorities do not always agree on questions of interpretation. The Commission’s 2024 review of the law records the protection authorities of three member states taking different views of the lawful basis for processing clinical-trial data, national guidance conflicting with the EDPB’s, and conflicting interpretations among authorities within the same member state.4 Moreover, these administrative layers are overlaid by national and European case law, with around 30 rulings by the Court of Justice of the European Union (CJEU) on the GDPR each year,5 rulings which may differ in emphasis or tone from those of national courts. A literature of practitioner commentary has grown around all of this, with one leading volume running to more than 1,300 pages.6 To know what this single law requires of a particular use, of a particular person’s data, in a particular jurisdiction, in a particular context, at a particular time, is to read across all of these layers at once. I’ve worked for almost a decade in the software of online communities and cannot remember a time of greater anxiety and confusion for the managers of online communities, many of which are not operated for a profit, than the passage of the GDPR. Even though I broadly support the GDPR’s aims, the kind of complex multi-layered law it typifies can have a significant cost on the communities, businesses and individuals it is (theoretically) intended to benefit, particularly if they are left unaided by the legal profession.

A similar situation holds for a number of legal regimes across most liberal democratic states. A few brief examples will serve to illustrate the point. In the United States (US), the Dodd-Frank Wall Street Reform and Consumer Protection Act, the legislative response to the 2008 financial crisis, has generated tens of thousands of pages of regulations.7 In her 2014 study, Roberta Romano described ‘the mind-boggling number of regulatory actions mandated by Dodd-Frank’ and identified the underlying pattern as the ‘Iron Law of Financial Regulation’: crisis-driven legislation, off-the-rack solutions, and a one-way regulatory ratchet that, over time, produces ‘an increasingly ineffective regulatory apparatus’.8 In England, four town and country planning statutes were consolidated in 1990 and have since been joined by six more, most recently the Planning and Infrastructure Act 2025, together totalling approximately 2,500 pages, alongside 8,000 pages of guidance, albeit it is now digitised.9 The overall effect has been described as ‘fragmented and confusing’, with ‘conflicting policy objectives’ and a legal framework that has ‘become more complex and confused, with fragmented legislation shaping differing aspects of local and national planning’.10 The Australian Law Reform Commission’s 2023 report on corporations and financial services legislation concluded that the existing framework ‘is no longer fit for purpose’,11 noting that the Corporations Act 2001 has nearly doubled in length since its enactment to more than four thousand pages.12 The report collected two decades of on-the-record epithets from Australia’s most senior judges concerning the most important piece of legislation in the Australian economy, including ‘exceptionally complex’, ‘labyrinthine’, ‘tortuous’ and ‘porridge’.13 These examples are far from exhaustive, and no doubt may be ameliorated by the detail of individual cases and concerns. Nevertheless, when the law of modern democratic states is taken in the aggregate, they are indicative of the overall trend toward complexity.

As well as growing more complex, laws have also proliferated in number. In 2024 the US Federal Register ran to 106,109 pages and 3,248 new rules.14 In 2025 the Trump Administration issued Executive Order 14192, which required federal agencies to repeal ten existing regulations for every new one proposed.15 Nevertheless, by the end of 2025 there were 2,441 new rules in the Federal Register.16 Indeed, during President Trump’s first term, Executive Order 13771 had already required two existing regulations to be repealed for every new one,17 yet the Government Accountability Office later concluded it had little impact.18 Between 2019 and 2024 the EU passed roughly 13,000 acts: 515 ordinary legislative acts, 2,431 other legislative acts, 954 delegated acts, 5,713 implementing acts and 3,442 others.19 The EU has recently layered the Digital Services Act, the Digital Markets Act, the Data Act, the Cyber Resilience Act, and the Artificial Intelligence Act onto a digital services regulatory environment that already included the GDPR. In the UK, the statute book stands at approximately 50 million words, with around 100,000 added or changed every month.20 Writing in 2007, Lord Bingham warned that ‘the sheer volume of current legislation raises serious problems of accessibility, despite the internet’, a consequence of ‘the legislative hyperactivity which appears to have become a permanent feature of our governance’.21 The trend has only deepened since. The annual count of new laws in the UK grew from 1,200 pages in 1960 to 2,700 pages in 2010, with the average Act growing from 24 clauses to 49.22 The UK Parliament added 1,800 pages of new primary legislation in 2024 and 2,500 in 2025. The Finance Act 2024 alone ran to 337 pages, more than twice its 2010 predecessor.23

The same ratchet of prolixity is occurring in courts as well. My first job after graduating law school in 2010 was working as judicial associate to Justice Christopher Carr in the appeal of The Bell Group Ltd (in liq) v Westpac Banking Corporation.24 The litigation started in 1995. The trial consumed 404 days of hearings during which 86,340 documents were tendered into evidence. The 2,643 page trial judgment was finally handed down in 2008.25 Justices Carr, Lee and Drummond were brought out of retirement (from the Federal Court of Australia), each were given an associate, and an entire courtroom was set aside just to hear a single appeal involving scores of lawyers and multiple senior barristers. It took me a month of singular focus to read and understand the trial judgment and many more months to read the tens of thousands of pages of written submissions. The appeal itself ran for over two years, and the written judgments on appeal ran to over a thousand pages of purely legal analysis.26 While that case is an outlier, it is emblematic of the direction in which the judicial process is headed in the face of the growing volume and complexity of laws and their disposition with sophisticated parties. Indeed, despite its Dickensian nature, the fact that the Bell case largely concerned events in the 1980s and 1990s meant that the Corporations Act 2001, which I previously noted has been judicially described as ‘tortuous’ and ‘porridge’, didn’t apply (at least not directly). The same expansion of judicial workload, judgment length, and procedural complexity has been documented by senior judges across common law jurisdictions. Bingham warned that ‘the length, complexity and sometimes prolixity of modern common law judgments … raise problems of their own’.27 Richard Posner charged that some judges complexify and mystify the practice of judging rather than confront ‘the growing complexity of the activities that give rise to the cases they must decide’.28 Nevertheless, I would not place too much blame at the feet of judges who have to regularly wrestle with the real and growing complexity of the law I’ve been laying out. Justice Neville Owen reflected on page 2565 of the Bell trial judgment that ‘from time to time during the last five years I felt as if I were confined to an oubliette’,29 a feeling to which many judges will be able to relate.

1.2 Complexity harms the rule of law

Lawyers have long been concerned by the growing unmanageability of modern law. Part of that concern lies in the challenges it poses to the performance of legal work, but, perhaps more fundamentally, in the risk it poses to the rule of law itself. In Bingham’s formulation, the rule of law requires that ‘all persons and authorities within the state, whether public or private, should be bound by and entitled to the benefit of laws publicly and prospectively promulgated and publicly administered in the courts’.30 That principle makes several demands, three of which are pertinent here. First, that the law be accessible, so that those bound by it can, without undue difficulty, find out what it is and base a course of action upon it. Second, that questions of legal right and liability be resolved by the application of the law rather than the exercise of discretion. And third, that the conduct of those who administer the law correspond to the law as announced.31 When regulation outruns the capacity of an enterprise, a lawyer, or even a judge, to know what compliance requires, then the law is no longer accessible to those it binds. When a regulator is unable to properly manage its own enforcement framework, the conduct of those who administer the law no longer corresponds to the law as announced. What results is not a softer rule, but a discretionary triage shaped by resources, attention, the political moment, or chance.32

A 2013 review into the causes of complex legislation in the UK summarised the risk it poses to the rule of law as follows: ‘Excessive complexity hinders economic activity, creating burdens for individuals, businesses and communities. It obstructs good government. It undermines the rule of law’.33 Daniel Greenberg, the current UK Parliamentary Commissioner for Standards, has previously suggested that ‘[d]angerous legislative trends are emerging that threaten the effective protection of the rule of law’.34 France’s Conseil d’ÉtatCouncil of State reached the same conclusion in successive annual studies of the state of French law. In 2006 it warned that ‘la complexité croissante des normes menace l’État de droitthe growing complexity of norms threatens the rule of law’, and in 2016, surveying ten years of reform attempts since, it found that ‘puisque le constat global est celui d’une dégradation de la qualité du droitthe overall finding is one of a degradation of the quality of law’.35 In a 2025 report, the Conseil summarised the concern as follows: ‘le droit s’est densifié au détriment de l’efficacité de l’action publique. Le paysage juridique est devenu progressivement plus lourd, moins lisible.law has thickened to the detriment of effective public action. The legal landscape has grown progressively heavier, less legible’.36 One of the more evocative expressions of the concern has been given by Hans-Jürgen Papier, the former President of the Federal Constitutional Court of Germany, in his 2019 book Die Warnung: Wie der Rechtsstaat ausgehöhlt wird:

Wie Mehltau hat sich die Überregulierung auf die Republik gelegt – wo das Recht doch eigentlich nur Sicherheit für kreative Initiativen und Aktivitäten bieten sollte. Mehr Gesetze bedeuten nicht automatisch mehr Recht und schon gar nicht mehr Gerechtigkeit.Like mildew, over-regulation has settled on the Republic – where law was only ever meant to provide security for creative initiatives and activities. More laws do not automatically mean more law, and certainly not more justice. 37

The risk that law’s expanding empire poses to the integrity of its rule is real and has been recognised as such by the leading lawyers, judges and officials in most liberal democratic jurisdictions for some time now.

To see this reality in action, we can turn back to the GDPR. A year after the GDPR took effect, researchers found that only about half of European small and medium enterprises believed they were fully compliant with it.38 The researchers noted that even that confidence weakened when specific compliance questions were asked. A contemporaneous survey of 293 UK SMEs found that only 10 per cent believed their organisations were fully compliant with the GDPR.39 Seven years on, the picture is no better. Data protection complaints to the UK Information Commissioner’s Office (ICO) rose in 2023–24 and 2024–25, reaching 42,315 in 2024–25 alongside 12,412 personal data breaches reported in the same year.40 The Office itself was systematically unable to keep pace with that volume of work. An analysis of the ICO’s data by David Erdos found that 70 per cent of data protection complaints in 2024–25 received no response from the Office within the required three-month period (up from 15 per cent a year earlier), only 3 per cent of reported data breaches resulted in any investigation, the number of UK GDPR investigations undertaken by the ICO fell by roughly 85 per cent in a single year (from 285 to 43), and the Commissioner issued no UK GDPR enforcement notices whatsoever.41 Motivated partly by this parlous state of affairs with the regulator, the Data (Use and Access) Act 2025 offloaded frontline complaint handling to businesses and organisations.42 Nevertheless, in November 2025 more than 70 civil society organisations, academics and data protection experts jointly called on the House of Commons Science, Innovation and Technology Committee to open a parliamentary inquiry into what they described as a ‘collapse’ of ICO enforcement activity.43 This is a relatively clear case of a regulator being unable to properly manage its own complex enforcement framework. The fundamental issue here is not the regulator per se, it is the overly complex nature of the legal regime they are trying to regulate within the resource constraints regulators always face.

1.3 Law reform is insufficient

Some argue that this problematic reality of modern law arises from a failure to put human experience at the centre of the legal order. They argue that we need to redouble our efforts to render laws and judgments in clear and concise prose,44 and remember that the law ‘is the witness and external deposit of our moral life’.45 As James Allsop, former Chief Justice of the Federal Court of Australia, puts it:

If legislation is to be built on complex and interlocking definitions, or if doctrine is to be ordered minutely in the attempt to express exhaustively the minute reach and particular application of the underlying norm, there comes a point where the human character of the narrative fails, where its moral purpose is lost in a thicket of definitions, exceptions and inclusions. The vice is not just lack of clarity; that is bad enough. Worse, it is a loss of human context, a loss of the expression of the human purpose of the law.46

Others have focused on the relationship between the regulator, the regulations, and the regulated. Cynthia Giles, who led the US Environmental Protection Agency’s Office of Enforcement and Compliance Assurance from 2009 to 2017, described the underlying dynamic of non-compliance with US environmental regulations as follows:

But often there is no one actively deciding to violate, just a series of ill-advised choices, or failures to choose, that result in a serious violation. People who try to characterize this as either ‘companies want to do the right thing’ or ‘companies put profits above people’ create barriers to solutions by putting a moral frame around what is actually just a practical issue.47

In his September 2024 report on the future of European competitiveness, Mario Draghi found that more than 60 per cent of EU companies consider regulation an obstacle to investment, and 55 per cent of small and medium-sized enterprises flag regulatory and administrative burden as their single greatest challenge.48 In response, the European Commission adopted the Competitiveness Compass, a five-year programme that commits the Commission to an ‘unprecedented simplification effort’ of a 25 per cent cut in reporting and administrative burdens for companies generally and 35 per cent for small and medium enterprises by the end of the mandate.49

Nevertheless, the calls for plain language, new approaches to regulatory design, or efforts at regulatory reform have sounded for many decades now, and while they may have helped around the edges, I would suggest they have been demonstrably unsuccessful in stemming the rising tide. There are many reasons for this, ranging from the fact that reducing the overall volume of regulation has a diffuse constituency, whereas individual regulations often have a vocal and persistent constituency, to the fact that the modern world is simply more complex, and its regulatory demands are greater and increasing.50 There is insufficient space here to systematically analyse why repeated attempts at regulatory reform over a number of decades have failed to address this problem. Suffice it to say for our purposes here that that failure is shown in the brief evidence I’ve already presented, and becomes only clearer the more evidence is examined.51 It may be possible that some short term progress can be made around the edges. However, I would suggest that when we increase the aperture of our view, such efforts are dwarfed by the overall trajectory, which Jonathan Sumption, former Justice of the UK Supreme Court, summarised as follows:

The expanding empire of law is one of the most significant phenomena of our time. Until the nineteenth century, most human social interactions were governed by custom and convention. The law dealt with a very narrow range of human problems. It regulated title to property. It enforced contracts. It protected people’s lives, their persons, their liberty and their property against arbitrary injury. But that was about all. Today, law penetrates every corner of human life.52

I suggest that the legal community needs to face the reality that the trends of the past hundred years are not going to be addressed by updating approaches to regulatory design, improving the use of language, pursuing new deregulation initiatives or yet one more invocation of Oliver Wendell Holmes. I have been seeing and experiencing this reality since the time I first entered law school, as have many other lawyers. In the face of the evidence, the reasonable assumption is that the same trends will continue, if not intensify, further threatening the rule of law. I suggest that an entirely new approach is needed. The rest of this essay will outline an Artificial General Intelligence for Law (AGIL) that attempts to meet this reality on its own terms.

2. Law can be mapped

[E]ven skilled lawyers have felt that, though they know the law, there is much about law and its relations to other things that they cannot explain and do not fully understand. Like a man who can get from one point to another in a familiar town but cannot explain or show others how to do it, those who press for a definition need a map exhibiting clearly the relationships dimly felt to exist between the law they know and other things.53

2.1 Law is relational

To many, including some lawyers, the application of law to human action is primarily concerned with the interpretation and application of text. The law of governments, parliaments, courts and contracts are made of sections, paragraphs, clauses and reasons all expressed in carefully worded prose. Lawyers seem to spend much of their time interpreting, analysing and applying these various texts in yet more documents containing advice, research, opinions and decisions. It would seem to follow from this that the solution to law’s expanding empire will have something to do with AI grounded in text, like an LLM, or perhaps an LLM packaged as an ‘agent’. There is an element of truth in that thought, but it also occludes something important about the nature of law.

Law is relational, both in its daily practice and in much of its normative structure. From the day we’re born till the day we die we exist in myriad legal relations with people, governments, objects, homes, banks, schools and businesses. And all of those actors with which we have legal relations have legal relations with each other too. A bank lives under the continuous oversight of its prudential supervisor, who examines, approves and restricts its actions. A sale of goods may happen over emails, calls and handshakes, all of which may shape the contract and any of its legislatively mandated terms. A landlord’s power to recover possession, a tenant’s right to remain, and a council’s powers to regulate are the matter of notices, grounds, emails, documents and hearings, playing out over many weeks, months and years. The relations are horizontal and vertical, public and private, physical and virtual. They are the ‘law in action’ as Roscoe Pound put it.54 The law in books attempts to set the rules that govern those relations. However, ‘law’ properly conceived is just as much about the law in action as it is about the law in books. A student thinks the law is about texts and rules. A lawyer knows it is also about relationships. It is perhaps this relational, active, understanding of the law that is one of the thoughts, subconscious perhaps, underlying many lawyers’ instincts that AI, however packaged, however good it gets, could never entirely, or even mostly, replace them. That response is partly reactive, but it is also partly grounded in the experiential understanding that text-based media, however sophisticated, can never completely model or comprehend what the law is like in practice. The relational territory of legal practice is not reducible to a model or simulation of it. The question for legal AI is not whether it can replace that territory of relationships, but whether it can navigate it. Whether it can draw a useful map and produce useful directional guidance. As Alfred Korzybski put it, ‘A map is not the territory it represents, but, if correct, it has a similar structure to the territory, which accounts for its usefulness’.55

Natural language processing is necessary, but not sufficient, to draw such a map. The ability to process, synthesise, generate and reason over text at scale has real value to existing ways of practising law, as we’re already seeing in the rise of legal research, project management and contract management wrappers around LLM-centric AI. These are significant developments. However, they are only the start of the journey. Just like you would not want to be in the passenger seat of a self-driving car being driven solely by ChatGPT, you would not want to be using a legal AI purely reliant on an LLM, no matter how well trained or refined it was. The insufficiency of an LLM to navigate the relationships of the law in action goes beyond fabricating cases, or getting the section number of a statute wrong. By itself, an LLM will always be structurally insufficient for the task of legal AI. Just like if you plugged some future, more advanced, version of ChatGPT into a regular car and asked it to drive you across town, an advanced LLM would still not have the appropriate interfaces, scoping, sensors, accuracy, error handling, redundancies and system design to achieve the goal. Nevertheless, also just like in the case of self-driving cars, neural networks will play a critical role in achieving the goal. The next significant step toward AGIL will be to build a system that can navigate the social and economic relationships that constitute the law in action by integrating the ability of LLMs to process and reason over text into operational systems grounded in a normative model of law. To give it its formal name, the next step will be neuro-symbolic legal intelligence, ‘neuro’ being short for neural network (eg LLMs), and ‘symbolic’ being short for symbolic AI, the field of AI concerned with the structural representation of knowledge. In the paragraphs that follow in this section, and in the subsequent section, I will sketch out the jurisprudence I think is relevant to a symbolic map that can ground a reasoning LLM, in such a way that AGIL can be achieved. For those already familiar with the works and ideas of jurisprudence and logic I mention, my accounting of them may seem superficial or selective. The point is not to engage in jurisprudential or philosophical analysis per se, it is to note what features of jurisprudence and logic may be relevant to a map of the law that may be used in combination with neural networks to achieve AGIL. For those not familiar with the works I mention, this brief sketch will also be insufficient, as it is merely a sketch. Nevertheless, understood correctly, it is a sketch pointing to the structure needed to build an AGIL.

To begin to create a symbolic structure of use to a neuro-symbolic AGIL, we can start by observing that the relationships which we’ve said form the substance of law are capable of being classified according to a scheme that arguably holds for any given law. In the early twentieth century Wesley Hohfeld published two investigations into the nature of legal relations in 191356 and 1917.57 Hohfeld started by noting the difference between the ‘physical and mental facts that call such relations into being’ and their legal disposition.58 The way we typically understand the relationships between people, and between people and things, is not the same as their legal understanding:

[P]hysical relations are wholly distinct from jural relations. The latter take significance from the law; and, since the purpose of the law is to regulate the conduct of human beings, all jural relations must, in order to be clear and direct in their meaning, be predicated of such human beings.59

The paradigmatic example of this is property rights. In ordinary speech, we identify property with a physical thing. We say that a car, a pen, or a house is ‘our property’. Legally, however, property is not the thing itself but a recognised relation between legal actors concerning the thing. Property is, in fact, various types of interests that organise relations among owners, tenants, purchasers, lenders, public authorities and others. The available types of property interest are restricted by what is known as the numerus claususclosed number principle, which influences legislative and judicial approaches across the civilian and common law traditions.60 Property therefore can be seen to involve two limbs. Certain combinations of facts constitute recognised interests, and those interests structure the legal relations among the relevant parties. Hohfeld generalises the second limb, legal relations among parties, by analysing possible legal positions into rights, privileges, powers and immunities, arranged as correlatives and opposites.61 Correlatives describe the same jural relation from the positions of its two parties, for example if A has a right against B then B has the corresponding duty, whereas opposites identify positions that cannot be held by the same person with respect to the same facts. His scheme is often represented in tabular form:

Jural Opposites

rightsprivilegepowerimmunity
no-rightsdutydisabilityliability

Jural Correlatives

rightprivilegepowerimmunity
dutyno-rightliabilitydisability

This scheme attempts to decompose the relations in any given legal discipline into elementary components that can be represented within a more general analytical system. Taking an archetypal example from English property law, suppose John owns Blackacre, Mary owns neighbouring Greenacre, and a deed grants a right of way over a path across Blackacre for the benefit of Greenacre. We could translate that into:

Correlative pairMaryJohn
Privilege / no-rightMary has a privilege to use the path.John has a no-right that Mary stay off the path.
Right / dutyMary has a right that John not substantially interfere with her use.John has a duty not to substantially interfere with Mary’s use.

What we might ordinarily call Mary’s “right of way” comprises more than one jural relation. Mary has both a privilege to use the path and a right against substantial interference.

Hohfeld spills considerable ink attempting to ground his scheme in the English common law as it stood at the beginning of the twentieth century, and elaborate the specific legal nuances of the concepts his scheme uses, such as ‘right’.62 Justifying or critiquing Hohfeld’s scheme, or its numerous interlocutors from the last hundred or so years, is not necessary here. Rather, the point is that to begin to sketch out a map that can be used by an AGIL to navigate the territory of law, fundamental ‘jural relations’, as Hohfeld calls them, are a good place to start as they provide a relatively clear path to operational experimentation with a generalisable structure in the context of real-world cases. To put it in practical terms, you can create a symbolic program that uses Hohfeld’s taxonomy of jural relations to break down a law into computational structures over which an LLM can reason.63 Indeed, it is important to keep in mind, as we march through over a hundred years of jurisprudence and logic, that the point of doing so is to lay the groundwork for such construction. When each new philosophy or discipline is mentioned, the question is not primarily whether any particular expression of the philosophy or discipline is more accurate or persuasive than another expression, albeit that question is still secondarily relevant. The primary question for this essay is the structural role that philosophy or discipline can play within a neuro-symbolic AGIL. In some cases it is important to choose a particular type of expression of a philosophy or discipline, not from a philosophical perspective, but from a practical one. In other cases all prior attempts at providing a grammar for analysis have failed, and one must experiment anew. Like all technology projects, the most important heuristic for any approach is whether it produces usable outputs from real-world inputs. In this sense the final analysis must be empirical.

Classifying jural relations is only the start of the story, the first broad lines on our map. We next need to observe that jural relations lie within a larger order that constitutes a legal system. Indeed, jural relations, however classified, can only have meaning, and, as such, can only be generalisable in the way required by AGIL, within a specific understanding of a legal system. As Hans Kelsen puts it:

Law is not, as it is sometimes said, a rule. It is a set of rules having the kind of unity we understand by a system. It is impossible to grasp the nature of law if we limit our attention to the single isolated rule. The relations which link together the particular rules of a legal order are also essential to the nature of law.64

Needless to say, how to think about the legal system within which such relations exist is contested. For the purposes of an AGIL, the nature of the project requires the conception of the system to be positivist, as opposed to one of natural law, legal realism, or legal interpretivism. We will see below that each of those other approaches to legal systems can inform the concept of the legal system used by an AGIL. However, an AGIL does need a positivist concept of law reflected in its structure so it can produce some form of legally recognisable output to hold utility as a legal reasoner in real-world contexts. Moreover, any AGIL needs a stable and consistent systems-philosophy to be able to process novel facts at scale while maintaining a verifiable level of quality in its outputs, as that is understood within a legal context. With some nuances, which we will explore below, an AGIL is a positivist project, albeit the specific type of positivism employed is more one of systems design than it is of analytical correctness. A system may be structured around a distinction between primary rules of obligation and secondary rules of recognition, change and adjudication,65 its sources,66 a hierarchy of norms,67 a classification of institutive, consequential and terminative rules,68 a system of shared plans,69 or a brand new positivist concept of a legal system. The heuristic a designer of an AGIL should use is the utility of the concept of the legal system to the symbolic system’s goals, and its success (or otherwise) in grounding the work of the neural networks it employs. All attempts at AGIL will employ a concept of a legal system, the question is how consciously that employment is executed, and what computational utility is produced by it.

In addition to a concept of a legal system, an AGIL needs to have a concept of the relations between different legal families. We should have perhaps started the history of ‘jural relations’ with Friedrich Carl von Savigny, one of the key figures of nineteenth century German jurisprudence, who used the phrase (Rechtsverhältnisse) in the sense we have used it in his analysis of Roman law.70 His account can be read as a historicist antecedent of an approach that Hohfeld later recast in analytical terms. One way to think about the attempt of an AGIL to define its fundamental legal system and basic taxonomy of legal relations is as an exercise in comparative law. Kelsen’s articulation of this approach is particularly apt because he presented his English-language General Theory of Law and State as an attempt to extend a theory originally formulated for civil-law countries so that it would also embrace the institutions of English and American law. As he puts it in the preface:

This theory, resulting from a comparative analysis of the different positive legal orders, furnishes the fundamental concepts by which the positive law of a definite legal community can be described. The subject matter of a general theory of law is the legal norms, their elements, their interrelation, the legal order as a whole, its structure, the relationship between different legal orders, and, finally, the unity of the law in the plurality of positive legal orders.71

Whatever flavour of positivism it employs, to earn its ‘general’ appellation an AGIL should have an understanding of law as a system of defined relational norms that can be seen to be reflected in the foundational jurisprudence of the world’s two major legal families. One might be tempted to call such a structure a ‘universal grammar’ of jurisprudence, loosely analogous to the linguistic hypothesis that a deep generative structure underlies the surface variety of human languages.72 Whether there really is such a grammar of law is a question we can leave open. Mauro Barberis observes that a general theory of law need not depend on this ‘strong’ conception of universality.73 It is enough that its concepts possess a ‘weak’ universality, which claims not that the same concepts are necessarily present in every legal order, but that they are susceptible to translation from one legal language into another.74 William Twining gives this thesis a Hohfeldian formulation, treating Hohfeld’s scheme as a case study in whether concepts ‘travel’.75 For Twining whether a concept travels well is not a metaphysical question but one answered by empirical and pragmatic tests: does it fit, does it work, and can it be used with roughly the same meaning in different settings?76 Twining understands Hohfeld’s scheme as an analytic vocabulary for decomposing and comparing legal discourse, not as a claim that each legal system must itself employ equivalent doctrinal terms. The lack of semantic equivalence does not defeat the normative scheme, albeit difficulty rendering the more basic modalities of ‘ought’, ‘may’ and ‘can’ may indicate that something is being lost in translation.77 The question of whether or not something is indeed lost is ultimately an empirical one. On this more practical footing, the systematic concepts jurists use to order their materials, such as right and duty, power and liability, are weakly universal not because every legal system must contain identical versions of them, but because they can be carried across the boundary from one system to the next if appropriately translated.78 Moreover, weak universality leaves room for the fact that some institution-specific concepts resist clean translation and behave more like proper names than like general concepts.79 In a computational system these are exceptions to a general typology, not barriers to implementation. Whether or not there is an ultimate metalanguage or metaphysics of transcendental legal universals, an AGIL need not rely on one. What it needs is the right system design.

To give a brief example, we can translate our previous example of Mary’s right of way over Blackacre into French law through the medium of Hohfeld’s correlative pairs. Suppose Jean owns fonds Aproperty A and Marie owns the neighbouring fonds Bproperty B . An acte notariénotarial instrument establishes a servitude de passageright of way over fonds Aproperty A for the benefit of fonds Bproperty B. Fonds Aproperty A is the fonds servantservient tenement and fonds Bproperty B the fonds dominantdominant tenement . We should note that the Code civilCivil Code provides that the owner of the fonds servantservient tenement may do nothing that tends to diminish the use of a servitude or to make its exercise more inconvenient.80 We can then express a similar table of relations as we did above:

Correlative pairMarieJean
Privilege / no-rightMarie has a privilege to pass over fonds Aproperty A within the scope of the servitude.Jean has a no-right against Marie’s passage within that scope.
Right / dutyMarie has a right that Jean not diminish the use of the servitude or make its exercise more inconvenient.Jean has a duty not to diminish the use of the servitude or make its exercise more inconvenient.

The English and French laws are not identical, but their relational structure is translatable. When each case is reduced to Hohfeld’s correlative pairs, we see that Marie/Mary has the privilege to use the path, John/Jean has no-right that she refrain, and Marie/Mary’s right against interference corresponds to John/Jean’s duty not to interfere. Of course, not all laws will be so straightforward, and neither does Hohfeld’s general scheme capture the nuances of even these simple examples. Indeed, it would be a mistake to attempt to fit all laws into a jurisprudential numerus claususclosed number of legal primitives. Rather, the point is that translation is possible at the conceptual level such that an AGIL can employ it where such translation is both necessary and practicable. As we will see below, this determination is sufficiently impacted by how one defines the scope and completeness of the operative legal system for any given analysis.

3. Law can be computed

3.1 Norms as normative propositions

Hohfeld’s classification can be read as an early attempt to render law in the form of a formal logic, but an AGIL needs to take that project considerably further. Hohfeld’s concern was a general scheme of jural relations grounded in the law of his day. An AGIL needs a logic that can scope, structure and verify the application of any given law to any given facts, and any assessment of whether the tests, standards and requirements of that law have been met. Twenty years after Hohfeld, and more squarely rooted in the field of formal logic, Jørgen Jørgensen posed a tricky problem for any attempt to apply logic to law.81 Inference, as classical logic understands it, is a relation between truth values. A conclusion is entailed because the premises could not be true while the conclusion is false. However imperative statements often found in legal expressions are neither true nor false. An imperative statement such as ‘shut the door’ is not a claim about the world, it is an injunction that the world be made a certain way. Nevertheless, we often reason with imperatives, and the reasoning appears to work. Keep your promises. This is a promise of yours. Therefore you should keep it. The inference seems to be sound, and yet its imperative premise bears no truth value. Either the inference is not valid or validity is not purely a question of truth. Jørgensen suggested that this dilemma could be resolved by holding that every imperative carries an ‘indicative’ component, a desired state of affairs, that can be extracted and reasoned over. Alf Ross objected to this response, arguing that focusing on the indicative component of an imperative formalises a description rather than a demand, and the force of the statement is lost.82 Georg Henrik von Wright suggested another answer to Jørgensen with a formal logic of obligation, permission and prohibition called ‘Deontic Logic’.83 Von Wright treated obligation, permission and prohibition as a single interdefinable family in a similar way to how necessity, possibility and impossibility form a family in alethic modal logic. Prohibition, what one is obliged not to do, is balanced against permission, what one is not obliged not to do. Hohfeld’s relations are a static list of opposites and correlatives, whereas von Wright’s are a calculus. A deontic calculus promises to reason over relations such that there is a possibility that deductive conclusions follow from a body of norms and facts.

Carlos E Alchourrón took this a step further by distinguishing deontic logic from normative logic. Whereas deontic logic concerns the logical properties and relations of norms without treating the norms themselves as true or false, normative logic concerns normative propositions, which do bear truth values.84 A normative proposition is a descriptive claim about a norm. It is a claim that the norm is part of a given system of norms. The proposition is true if the norm is indeed a coherent part of that body of norms. If we return to John and Mary and recall that John executed a deed granting Mary her right of way over Blackacre. Let E denote the applicable rules of English law, let d denote the deed, let p mean that Mary uses the path within the scope of the grant, and let ⊢ mean ‘follows from’. We can then express the ‘privilege’ component of the right of way as:

E, d ⊢ Pp

This is the normative proposition that Mary’s privilege to use the path follows from English law together with the deed. Note that its truth depends on whether that conclusion follows, not on whether Mary ever uses the path. In Alchourrón’s terms, this is a strong permission. It differs from:

E, d ⊬ O¬p

which says only that no norm from English law and the deed requires Mary not to use the path. That is weak permission. In other words it is the absence of a prohibition. I’d further note that using such a calculus significantly increases our range of expression. For example a right of way may arise in English law without a deed. Mary and the previous owners of Greenacre may have used the path for many years, such that there is a prescriptive easement, which we can denote with r. The expression then becomes:

E, r ⊢ Pp

These are just brief examples to give a sense of what we mean by a “normative proposition” and what expressive possibilities it has in representing legal scenarios.

One of the earliest attempts to apply formal logic to legal analysis as a lawyer would understand it was a paper by Layman Allen in 1957 where he made a case for the role of logic in the drafting and interpretation of legal instruments.85 Even though his paper came after von Wright’s on deontic logic, the academic streams had not intermingled at that point, and Allen employed ‘classic’ formal logic.86 Moreover, Allen’s concern was more syntactic than philosophical. A provision held together by conditionals and connectives may admit of several readings depending on the scope of an ‘unless’, or what an ‘or’ attaches to, and at times even the drafter cannot tell which reading was enacted. ‘Normalisation’ in formal logic makes the structure visible, so that a sentence with four possible readings shows those four readings. Interestingly, Allen went as far as computing a truth table to establish that two independent sections of the US Internal Revenue Code (IRC) were functionally equivalent:

This should persuade even the most dubious that 13.6 is equivalent to section 74 and the relevant portion of section 117 as they are now written in the Internal Revenue Code. In both, every imaginable combination of relevant facts leads to the same conclusion about whether the amount is to be included in gross income.87

Ronald Stamper’s LEGOL project took that idea further, toward something closer to a specification language for legislation.88 In a similar vein, albeit more abstracted, Allen and Charles Saxon developed A-HOHFELD, a pseudo-computational rendering of Hohfeldian relations.89 These efforts don’t have direct utility to an AGIL insofar as their application of logic to law is not operationalisable in general legal reasoning, albeit the general approach of comparing formalisations does, for example in testing different drafts of a provision(s). A different approach to the application of formal logic to legal concepts was taken by Stig and Helle Kanger, and subsequently refined by Lars Lindahl. Stig Kanger combined an impersonal deontic operator for what shall be the case with an action operator for an agent seeing to it that something is the case.90 Stig and Helle Kanger then used that calculus to construct 26 mutually exclusive and jointly exhaustive two-agent atomic types of rights.91 Lindahl refined the analysis by distinguishing passivity from action and counteraction, and constructed seven one-agent types, 35 individualistic two-agent types and 127 collectivistic two-agent types.92 We can get a sense of what this means by returning to John and Mary. Let’s imagine there is a gate across the path, raising the question of how its operation affects both Mary and John’s positions. Lindahl’s types allow us to express some of the possible permissions and obligations vis-à-vis the gate as follows:

Normative positionAllocation of permissions and obligations for the gate
T₁/T₁ (separate freedom)Mary and John may each open, close or remain passive.
T₁/T₆ (asymmetric freedom)Mary may open, close or remain passive. John must remain passive.
T₃/T₆ (required choice)Mary must either open or close, but may choose which. John must remain passive.
T₅/T₂ (required opening)Mary must open. John may open or remain passive, but must not close.
T₂/T₄ (opposing permissions)Mary may open or remain passive, but must not close. John may close or remain passive, but must not open.
Collectivistic position (shared obligation)At least one of them must open, although neither is individually required to do so.

The point of enumerating such possibilities is that it illuminates the ways a rule(s) could operate in this scenario, and how the action(s) of the relevant agents might indicate compliance or non-compliance with such a rule(s). Such an enumeration of different positions doesn’t show Hohfeld’s simpler typology to be wrong. The point is that Hohfeld identified more elementary jural relations, while the Kangers and Lindahl created a larger space of compound normative positions under stated logical assumptions which can operate in a computable calculus. Moreover, what matters here is not the lists of possible positions per se, but how they are generated. The key point is that once the relevant agents, acts and states of affairs have been specified, the chosen calculus can generate the space of logically possible positions.

3.2 Computing normative propositions

The first substantive attempts to develop legal logic into a form of computational law came in the 1970s and 80s. One noteworthy effort was LT McCarty’s analysis of his TAXMAN program in 1977, which modelled the taxation of corporate reorganisations.93 McCarty chose that focus for the program as he wanted it to be simple enough to model, while containing sufficiently lush and thorny concepts so he could test the limits of the experiment.94 Any formal model necessarily omits details that may matter in some contexts, so McCarty tried to push his model to its limits and use any unacceptable conclusions to expose its inadequacies.95 He found that many of the more mechanical statutory concepts could be represented without much difficulty. Section 368 of the IRC, for example, makes ‘control’ part of several definitions of reorganisation, while subsection (c) defines it by ownership of stock carrying at least 80 per cent of total voting power and at least 80 per cent of the total number of shares of all other classes. That is a threshold, applying to a single state of the world, with an algebraic formula at its core, and so it programs more or less directly.96 However, the judicial doctrines of continuity of interest, business purpose and step transaction were not so amenable. The courts had built those doctrines in opposition to the mechanical manipulation of the statutory rules, and, perhaps unsurprisingly, McCarty found they had a different structure from the logical program he had constructed.97 One explanation he offered for this limitation was that the programming language gave TAXMAN dynamic power and flexibility, while its representation of concepts remained statically equivalent to a higher-order predicate calculus.98 In later work, McCarty and NS Sridharan supplemented logical templates with a prototype-plus-deformation model.99 Concepts that did not admit a single logical template could be organised around concrete exemplars and transformations mapping one description into another, while a template could still supply an invariant component where one was available. Also of note from this period is Sergot, Kowalski and colleagues’ formalisation of the British Nationality Act 1981 in Prolog, a system that could derive a person’s citizenship from formalised statutory provisions once the facts of the case were supplied interactively.100 Defined concepts were reduced through the rule hierarchy until the program reached predicates not defined in the Act, ie those that had to be supplied from outside the program. For example, whether someone was ‘of good character’ could not be derived from the formalised Act itself.101 It either had to be supplied externally or treated as an assumption qualifying the program’s answer. The authors presented the program as an experiment in applying legislation mechanically to individual cases, not as a general simulation of legal reasoning. They also found that the Act’s provisions for abandoned infants required conclusions drawn by default to be withdrawn when contrary information later emerged, which the authors identified as non-monotonic reasoning.102 I briefly mention these examples to give a sense of the tenor of the projects of the period, which further ranged from product liability103 to German civil law.104

Richard Susskind reflected on these early developments in computational law in an article published in 1986 in which he advocates for an approach similar to that taken in this essay.105 Susskind found that most computational law projects made no reference to jurisprudence at all, and that where legal theory was mentioned it was not treated as a matter of central significance to the enterprise. With a handful of exceptions, the field had made marginal contributions to, rather than exploitations of, the jurisprudential resources available to it, and he suggested that this might be why no successful system yet existed.106 His response was that a project in computational law must be grounded in a jurisprudential understanding of a legal system, and that all such attempts are so grounded, whether they are cognisant of it or not:

[A]ll expert systems must embody a theory of structure and individuation of laws, a theory of legal norms, a theory of descriptive legal science, a theory of legal reasoning, a theory of logic and the law, and a theory of legal systems, as well as elements of a semantic theory, a sociology and a psychology of law (theories that must all themselves rest on more basic philosophical foundations). If this is so, it would seem prudent that the general theory of law implicit in expert systems should be explicitly articulated using (where appropriate) the relevant works of seasoned theoreticians of law.107

Susskind went on to enumerate eight problems that needed addressing before any such system would be capable of operating as an expert system in law, and considered that they would require decades of focused attention.108 Susskind was quite prescient, albeit time and the march of technology have changed the terms of his analysis. He concluded (in 1986) that intensive coverage of a small legal domain is preferable to superficial coverage of an extensive area, because ‘expert systems in law ought to be designed to replicate legal experts, the knowledge represented, therefore, necessarily being of a depth, richness and complexity normally possessed by such a human being’.109 The premise is essentially that depth is bought at the cost of breadth. That held in 1986 more as a matter of engineering than jurisprudence per se.110 Subsequent developments in knowledge graphs, structured models of legal concepts, LLMs to process text at scale, and many changes to programming languages and software systems, let a system take its depth from the legal sources themselves, which are now reachable at scale over a global internet with programmable interfaces, rather than from what a specialist can articulate about them. As such, generalisable depth becomes a question of system design rather than of expertise traditionally conceived.

Anne von der Lieth Gardner’s work in this period, which Susskind listed as one of the few to really engage with jurisprudence,111 is also an interesting ancestor to AGIL. In An Artificial Intelligence Approach to Legal Reasoning Gardner articulates a goal for computational law which applies directly to AGIL:

The overall objective is not a program that ‘solves’ legal problems by producing a single ‘correct’ analysis. Instead, the objective is to enable the program to recognize the issues a problem raises and to distinguish between those it has enough information to resolve and those on which competent human judgments might differ.112

Her charge against the burgeoning field of computational law of her time was that it had treated legal issues as though they were alike. Some programs attempted to weigh factors on hard questions, while others attempted some form of deduction, reverting to the user where deduction did not work.113 What she wanted was a program that could tell these scenarios apart, reaching an answer deductively where it could and concluding elsewhere that the human decision maker has room for choice, with the program providing the technically defensible analyses or factors and the human deciding among them.114 She took that division, between deduction and judgment, to be a permanent feature of the field rather than a concession to the state of the art, since programs will not be in a position to make that choice until we are satisfied that we have analysed fully, and encoded correctly, the concepts of wisdom and justice.115 She used HLA Hart’s concept of open texture as the limiting case between the poles of deduction and judgment.116 Her own solution to navigating this region was that the program should develop a set of heuristics to distinguish easy cases it can solve itself from hard cases it cannot resolve itself. As to the heuristics themselves, she suggests they be drawn from the use of the relevant predicates in both ordinary language and legal cases.117 This resulted in a classification of the region of open texture into predicates with a clear prototype case and many possible variations, definitions that look analytic but are in fact defeasible, and concepts so abstract that their range can be worked out only piecemeal.118 While her solution is not feasible, particularly at scale, Gardner’s description of the problems faced by an AGIL is one of the clearest.

3.3 Arguments can be computed

Back within the weeds of formal logic, the next relevant step came with Phan Minh Dung’s work on argumentation, which supplied a general way to represent what happens when reasons conflict and conclusions remain open to defeat.119 Prior research in AI and logic programming had produced several apparently different formalisms for non-monotonic reasoning and Dung showed that they could all be recast as forms of argumentation. His framework consists simply of a set of abstract arguments and a binary attack relation between them. An argument is acceptable relative to a set when every argument attacking it is itself attacked by a member of that set, and different semantics determine which admissible sets count as extensions. Let’s return to Jean and Marie to see this in action. Suppose the reason Jean installed the gate across the path was to contain some sheep and he gave Marie a key. On the information initially available, Jean advances argument A_J: Marie has a key and the delay to her progress along the path is negligible, so the gate does not make the exercise of the servitude more inconvenient. With no attacker, A_J is accepted. However new evidence then comes to light which suggests two more arguments. A_L states that the lock repeatedly jams and delays Marie by two to three minutes several times each day. A_M concludes that these recurrent delays render the exercise of the servitude more inconvenient. A_J and A_M attack one another, since they both can’t be true, while A_L attacks A_J’s premise that the delay is negligible. We can express this as:

A_J ↔ A_M A_L → A_J

Suppose the system now accepts A_L, which defeats A_J and thereby defends A_M. The earlier conclusion required by A_J is withdrawn and the contrary conclusion required by A_M accepted. This is an example of argumentation producing non-monotonic reasoning, ie a change in the outcome based on new information.

Dung’s argumentation framework has been brought into law in a few ways worth briefly listing. Henry Prakken and Giovanni Sartor built a dialectical model on it in which arguments are tested through attack and defence, conflicts are resolved by priorities among rules, and the priorities used to resolve one conflict may themselves be supported or defeated by further arguments.120 The ASPIC+ framework of Sanjay Modgil and Prakken gives Dung’s arguments an internal structure built from strict and defeasible rules, and separates three ways an argument can be attacked: rebutting its conclusion, undercutting the rule that licenses it, or undermining one of its premises.121 Thomas Gordon, Prakken and Douglas Walton’s Carneades model added proof standards, and the allocation of the burden of proof, which decide which party must establish what before a conclusion will hold.122 Trevor Bench-Capon’s value-based argumentation framework rests on the recognition that judges and legislators weigh not bare arguments but the values those arguments advance, so that whether an attack succeeds can depend on how a given audience ranks the competing values,123 a thought which Sartor developed further into a logic of teleological reasoning and proportionality for the weighing of rights and values.124 Bench-Capon and Katie Atkinson applied a factor and value based analysis to the ANGELIC methodology, which organises legal domain knowledge into linked issues, factors, acceptance conditions, sources, values and questions for reasoning about cases.125 Guido Governatori’s framework combined defeasible logic for representing exceptions and priorities between conflicting rules with a deontic logic of violations in which reparation chains encoded contrary-to-duty obligations.126 Governatori tested his framework against the 2012 Australian Telecommunications Consumer Protection Code in a real-world pilot.127 There are a number of other authors and papers that have attempted to extend the reach of defeasible or argumentative logic into law that could be mentioned here, some more theoretical, and some more practical. The point here is not to be exhaustive, critique, or add to this list. Rather, as we will see more fully in the next section, the goal is to find a coherent way to usefully employ this developed formal logic of law within a coherent approach to general legal intelligence.

In the last few years a further wave of research has set out to combine this work in formal logic with the power of LLMs, ie to engage in experiments with neuro-symbolic computational law. In briefly sketching the recent neuro-symbolic efforts I will also list the specific models used in each, to give a sense of where they might sit vis-à-vis the models available at the time of this essay. Indeed, it is perhaps misleading to refer to “LLMs” as a single thing, as their capabilities can, and will continue to, vary significantly. Some recent work uses LLMs to derive principles from legal texts. Morgan Gray and colleagues used GPT-4 turbo and Llama-3.1 to read a set of cases and propose the factors that produce their outcomes.128 In a companion study the same researchers used GPT-4o, Claude 3.5 Sonnet and Llama-3 to gauge the magnitude of factors set by experts, specifically the degree to which each favours one side of a case, and generated arguments that turn on those relative strengths.129 I note here that I will analyse these two case-factor studies in some depth below. Hannes Westermann and colleagues compared a number of models, including GPT-4o, Llama-3.3, Claude 3.5 Sonnet and Gemini 2.0 Flash in mapping the text of adjudicatory decisions onto a fixed checklist of legal criteria, recording for each criterion whether the decision-maker found it satisfied and explaining why.130 That work rests somewhat on Jaromír Šavelka and Kevin Ashley’s finding that GPT-4 and GPT-3.5 annotate legal texts tolerably well with no task-specific training.131 Other recent work uses formal structures to organise or constrain what an LLM does. Gabriel Freedman and colleagues used multiple models, including GPT-3.5, GPT-4o and Llama-3, to build a formal argumentation framework, so that the verification of a claim can be explained and contested rather than merely asserted.132 Manuj Kant and colleagues had OpenAI’s o1-preview translate a simplified hospital cash-benefit policy into Prolog rules and run them to adjudicate claims, with the o1-preview proving more reliable than the ‘non-reasoning’ GPT-4o,133 and a follow-up handled the harder coverage questions under a US student health plan by giving a range of reasoning models, among them OpenAI o1 and DeepSeek R1, an expert-curated framework to work within.134 Jiseong Chung and colleagues aligned a policy graph drawn from the articles of the GDPR with a context graph extracted from the scenario in hand, and routed the analysis through a deterministic compliance gate that traverses the cross-references and hands a judge model, GPT-4.1 by default, only a constrained question.135 Others have paired an LLM with the Z3 satisfiability modulo theories solver to test enforcement actions under Taiwan’s Insurance Act for consistency,136 or post-trained open models, among them Qwen3 and Llama-3.1, against a knowledge graph of decided cases structured on the common pattern of issue, rule, analysis and conclusion.137 A similar approach has been commercialised by Norm Ai, the venture that grew out of John Nay’s Law Informs Code, which renders regulations as programs that wrap a language model for financial services compliance.138 While each of these neuro-symbolic projects is valuable and interesting in its own right, they are all experiments on specific laws, actors, capabilities or techniques, and are not situated within a broader concept of a legal system such that one could say they are engaged in the project of AGIL. Nay’s broader outlook in Law Informs Code and his associated project of Norm Ai is perhaps an exception to this and I will address his ideas later in this essay. But we first must leave the firm ground which we’ve been traversing for some more challenging terrain.

4. Law can be messy

There was a man in our town
and he was wondrous wise:
he jumped into a bramble bush
and scratched out both his eyes—
and when he saw that he was blind,
with all his might and main,
he jumped into another one
and scratched them in again.139

No map captures everything its territory contains. There will always be thickets of unforeseen, contested, unclear and ambiguous cases that legal practice must navigate. Karl Llewellyn took The Bramble Bush, the title of his primer for law students, from the eponymous nursery rhyme that evokes this thorny reality of the law.140 The law in action is messy, complex, overlapping and even contradictory. The map has unfinished lines, contested borders, and sometimes ends with the note ‘here be monsters’. If we render the modern history of jurisprudence and formal logic as a linear progression of increasing clarity, reaching its apogee in a defeasible, non-monotonic, deontic calculus encoded into machine readable form and augmented by an LLM, we miss something crucial, something which, again rightly, feeds into some lawyers’ scepticism about AI. The law of regulators, courtrooms, contracts and council chambers does not arrange itself for the convenience of its systematisers. Somewhat counterintuitively, the law in action is also where LLMs, for all their power, are most risky. Ask a general-purpose model a precise question of law and its answer can be inconsistent with its prompt, its own training, or with the law or facts.141 Lawyers have already been fined for citing cases a model had invented.142 One attempt at a remedy is to connect the model to a trusted legal database, ie using retrieval augmented generation, but that does not really address the fundamental problems.143 The issue is the inherent statistical nature of LLMs, and how the billions of parameters within their matrices are reached, or in other words how they are ‘trained’ and ‘refined’. Common methods of training and refinement of LLMs reward a confident guess over an honest ‘I don’t know’.144 Models do not reliably communicate their uncertainty and, in the context of ungrounded legal question-answering, have been shown to express misplaced confidence in hallucinated answers and accept false premises.145 Moreover, a model’s own account of its reasoning cannot be taken for granted. Models omit factors that influence their answer such that their so-called ‘chain of thought’ is not representative of how they reach conclusions, and not equivalent to an audit trail.146 In a recent study of legal entailment, models that were given explicit formal representations of contract clauses and asked to reason with them reported conclusions that diverged from what the symbolic execution of those same representations produced.147 Models were found to reason informally and then present the output as if it were derived from the logic, a behaviour the authors call ‘scope laundering’.148 Moreover, the scoping, accuracy and auditability problems become acute at any kind of scale. Evaluations of long-context models have found substantial degradation as contexts grow, especially where locating the relevant material requires semantic inference rather than literal matching.149 All of these weaknesses play a role in LLM’s inability to deal, when operating alone, with the positivist black letter law. However, they become even more acute as we start to move beyond black letter law. Inexactness at scale about legal reasoning that can be scoped, validated, and audited by formal logic is one thing. Inexactness at scale about legal reasoning that is not so amenable to formalisation is a non-starter, even if models become considerably better than they are now. This is especially true if the inexactness resides in a reversion to a mean in an area of law where cases sometimes allow for exceptions to that mean. In any event, before we can start to consider how symbolic structures and neural networks might deal with law’s brambles, we need to first describe what these ‘brambles’ I’ve been referring to actually are.

4.1 How law handles mess

In the common law tradition, the legal brambles are partially the domain of the law of equity, which rose as a check on, extension to, and salve for black letter law. The term originally comes from Aristotle, who described the nature of epieikeiaequity as:

a correction of law in the respect in which it is deficient because of its being general. For this is the cause also of the fact that all things are not in accord with law: it is impossible to set down a law in some matters, so that one must have recourse to a specific decree instead.150

In other words, any system of rules of general application will produce some outcomes that nobody wants, and the law needs a way of dealing with that. Christopher St German’s Doctor and Student gave the first theoretical justification for equitable intervention in English law, drawing on Aristotle’s concept of epieikeiaequity as developed by Jean Gerson and Thomas Aquinas.151 When Lord Ellesmere came to give equity its canonical statement from the Chancery bench in the Earl of Oxford’s Case, he was repeating St German almost verbatim:

The Cause why there is a Chancery is, for that Mens Actions are so divers and infinite, That it is impossible to make any general Law which may aptly meet with every particular Act, and not fail in some Circumstances.152

What Lord Ellesmere described in 1615 is now a body of distinct doctrines, including equitable estoppel, unconscionable conduct, fiduciary obligation, equitable compensation, relief against forfeiture, constructive trusts, and even equitable fraud.153 What runs through equity’s doctrines is the posture that strict legal entitlement coexists with substantive standards of conscience, duty, and trust, and that the legal system must reserve to itself the means to respect them when the rules alone would not. The civilian tradition holds a similar posture, expressed through the principle of good faith and codified as a deliberately open general clause.154 The German Civil Code provides that ‘An obligor has a duty to perform according to the requirements of good faith, taking customary practice into consideration’,155 and from that sentence its courts have built whole doctrines the code never spelled out.156 However expressed, the analogous postures share a genealogy, seen more clearly through their antecedents, such as bona fidesgood faith in Roman law:

It was Quintus Scaevola, the pontifex maximus, who used to attach the greatest importance to all questions of arbitration to which the formula was appended ‘as good faith requires’; and he held that the expression ‘good faith’ had a very extensive application, for it was employed in trusteeships and partnerships, in trusts and commissions, in buying and selling, in hiring and letting - in a word, in all the transactions on which the social relations of daily life depend; in these, he said, it required a judge of great ability to decide the extent of each individual’s obligation to the other, especially when counter-claims were admissible in most cases.157

Modern legislatures add open standards to rules for the same reason, and courts often still rely on principles of equity and its civilian analogues to apply such standards, a good example being the interpretation of corporate statutory duties through the fiduciary doctrines from which they were drawn.158 Indeed, in the vocabulary of modern jurisprudence, dominated as it is by statute law, equity’s precepts are best conceived as ‘standards’.159 Where a rule fixes its content in advance, a standard leaves that content to be settled afterwards, in the light of the case at hand, as a matter of ‘judgment’. These standards are the part of law that most resists the symbolic treatment we have been describing, with their content broad, loosely defined and highly contextual, drawing on precedent, practice, policy, principle, and perhaps even moral judgment.

Legal realism developed an analogous response to these legal brambles, one that was more empirical and consequentialist in outlook. Though Anglo-Saxon lawyers tend to think of it as American, legal realism was a broad movement that also ran through European and Scandinavian jurisprudence in the first half of the twentieth century. Each strand tried, in its own way, to put ‘legal science’ on an empirical footing, and each set itself against the formalist assumption that law is a closed system of concepts from which decisions can simply be deduced. Part of that revolt was a new frankness about how judges actually make law, a tone present in the likes of Benjamin Cardozo,160 and perhaps most famously Oliver Wendell Holmes, who wrote that ‘the prophecies of what the courts will do in fact, and nothing more pretentious, are what I mean by the law’.161 As we’ve already noted, Roscoe Pound supplied the distinction between the ‘law in books’ and the ‘law in action’. His sociological approach is perhaps distinct from the American realists proper, and drew on European sources such as Eugen Ehrlich, whose idea of ‘living law’ located a society’s real rules in its everyday customary practice.162 Ehrlich belonged to a ‘continental’ approach that ran in parallel to the Americans. In France, François Gény’s libre recherche scientifiquefree scientific research gave the jurist licence to reason from social and economic facts wherever the written laws were insufficient.163 In Germany, the Freirechtsschulefree law movement of Kantorowicz treated judging as openly law-making,164 while Heck’s Interessenjurisprudenzjurisprudence of interests suggested judges weigh the interests behind a statute rather than deduce answers from its concepts.165 In America the same approach targeted the rules themselves. Llewellyn was a leading rule sceptic, and he argued that the ‘paper rules’ of the books do not by themselves decide cases.166 Conversely, Jerome Frank was skeptical about facts. He argued that the facts on which a case turns are not simply given to the court but are contested reconstructions produced through the process of litigation.167 He argued that legal certainty was a myth, sustained by professional habit and emotional need.168 The Scandinavian realists went even further. Axel Hägerström and his ilk treated positivist rights, duties and validity not as real things in the world but as a kind of metaphysical magic, in an approach that some have described as ‘legal nihilism’:

there are no legal concepts but only the use of words to express various feelings and interests among human beings with respect to what the proper behaviour is, evoking the appropriate response that can be enforced, if necessary, by means of coercive sanctions.169

We should be clear that most of these legal realists were not claiming that judging is inherently lawless, random or mystical. One could say that the median realist held that judicial decisions are largely predictable from the facts and from recurring ‘situation-types’ rather than from the rules themselves.170 The Critical Legal Studies movement of the 1970s and 1980s took this realist thesis in a more political direction in the US,171 and figures such as Pierre Bourdieu made a similar move in Europe.172 Its more consequentialist strand, meanwhile, flowed into what came to be called ‘law and economics’, particularly in the work of Richard Posner.173 The strong realist claim, that rules do little or no real work in deciding cases, does not need to be settled at this stage. At this stage I would simply ask a strong realist, of any type, to withhold their analysis until the outcomes of AGIL can be assessed, politically, economically or otherwise. In any event, the softer claim is perhaps the more common ground, namely that the formal picture of the law in books is an incomplete account of how law is created and applied, and that any system, human or computational, that operates as if it were will fail to be a good legal reasoner.

If we bring our appraisal of how the law deals with the messiness of its role in the world back into jurisprudence proper, we can observe that, as a general proposition, rules cannot be understood as having lexical force in the way a formal logic, however advanced, requires. As HLA Hart put it, many rules have an ‘open texture’,174 a penumbra of cases where their general terms do not plainly decide the question. Hart gives the example of a rule prohibiting the use of ‘vehicles’ in a public park. The rule clearly restricts the use of a car, but does not so clearly restrict the use of a bicycle or a pair of roller skates.175 The doubt, in each penumbral case, is not about the word(s) per se but about the relationships behind them: between rider and public, between vehicle and purpose, between the statute’s drafters and the world they sought to regulate. Law is made of written rules, but its application is not a question of lexical resolution, because the relational reality on which rules are brought to bear is richer than any positivist logic can exhaust. The concept of ‘open texture’ likely came to Hart from Friedrich Waismann — a close associate of Ludwig Wittgenstein — whose classes Hart attended in Oxford.176 As his thinking progressed, Wittgenstein notoriously broke with the logical positivists, arguing that the indeterminacy of ordinary language is not a failure to be corrected but may be what the use of language to communicate requires.177 The concepts we apply to the world such as game, number, or chair have no crisp defining conditions, rather they are held together by a ‘family resemblance’ of overlapping similarities, with no single feature uniting them all.178 The meaning of a word is not its definition but its use in language, and no rule, strictly, determines its own application, since every rule can be read in ways that accord with or diverge from any given case.179 The text of a rule cannot, by itself, decide whether a bicycle is a vehicle, because the concept ‘vehicle’ has no crisp centre to match against.

4.2 How systems handle mess

So how do we deal with these loose standards of equity and good faith, the realities of judgment, and the linguistic penumbra scattered across the legal landscape in an AGIL? When faced with the complexity of both nature and human societies, cartographers do not try to draw every coastline smoothly, every contour exactly, or every border with bright lines. They develop conventions for marking what cannot be precisely shown. There are legends of symbols, hatching for unsurveyed areas, or dashed lines for boundaries in dispute. In 1971 Alchourrón and Eugenio Bulygin published Normative Systems, which argued that the jurisprudence we’ve been describing has incorrectly framed the question of the completeness of a legal system.180 According to Alchourrón and Bulygin, positivist legal dogmatics and legal realism alike inherited an Aristotelian conception of system and tended to frame completeness at the level of the legal order as a whole, embracing all valid norms.181 Following this reasoning, epieikeiaequity is only a problem in the way we’ve been describing if we take this Aristotelian view of a normative system. Instead, they argue for an approach to a legal system based on Alfred Tarski’s concepts of deductive and axiomatic systems,182 situated within the relational reality of a legal system.183 We will need to spend some time breaking those claims down in order to see how they potentially allow an AGIL to handle law’s brambles.

Alchourrón and Bulygin note that there is the set of selected legal sentences or provisions in question for any given analysis, rarely the whole law of a jurisdiction and ordinarily a subset addressed to a particular problem.184 That range is sorted into potential cases by fixing a finite set of features of fact the analysis treats as legally relevant, the Universe of Properties, whose combinations generate the Universe of Cases.185 And there is the range of answers the rules could give to such a situation, the Universe of Solutions, typically the deontic qualifications of the actions in play, such as obligation, prohibition, or permission.186 Let’s turn back to Marie and Jean, and Jean’s new gate for his sheep, to see how this might work. Suppose that the norms selected for the analysis are: the acte notariénotarial instrument, Jean's right as owner to enclose his land under Article 647 of the Code civilCivil Code, and his duty not to diminish the use of the servitude or make its exercise more inconvenient as established by Article 701. Let’s also suppose that three facts about the gate form the Universe of Properties: whether the gate is locked, whether Marie has a key, and whether the gate causes her delay, each of which has two relevant states.

Property+-
The gate is lockedlockedunlocked
Marie holds a keykeyno key
Marie is delayedsubstantialtrivial

These properties produce eight different possible cases, or in other words our Universe of Cases. Furthermore, Jean’s installation of the gate allows for three possible solutions, or our Universe of Solutions: Jean may not keep the gate (‘prohibited’), Jean may keep the gate or not (‘facultative’), and, we must also add, Jean must keep the gate (‘obligatory’). Given the norms we’ve selected, the cases and solutions look like this:

CaseLockedKeyDelaySolution
1+++prohibited
2-++prohibited
3+-+prohibited
4--+prohibited
5++-facultative
6-+-facultative
7+--prohibited
8---facultative

Note that the third solution (obligatory) doesn’t appear, as it doesn’t result from the norms and cases we’ve selected, but it is still within the Universe of Solutions even if it is not used. We’d only need to add a clause to the acte requiring Jean to exercise livestock control for the ‘obligatory’ solution to appear. Indeed, norms, cases and solutions all have to be specified for the given analysis before ‘completeness’ can be assessed. A different specification of any one of them can change whether the system leaves a given case without a complete answer. Alchourrón and Bulygin put the point in these terms:

This implies that any discussion of the problem of gaps requires, as a preliminary step, the determination of the range of each of the three terms of the relation. It is the omission of this necessary step that has been responsible for the failure of legal philosophers to deal successfully with this problem.187

When norms, cases and solutions are fixed, different types of legal brambles can be identified as gaps in that system. A normative gap arises when the selected norms leave one kind of situation in the Universe of Cases without a full answer. The norms may be wholly silent or may determine only part of what is permitted, required or prohibited.188 For example, if we take Article 647 of the Code civilCivil Code out of the analysis above, cases 5, 6 and 8 are no longer connected with any solution by a norm. A gap of knowledge arises when we lack the factual information needed to decide which kind of case we are dealing with. For example, if we’re unsure whether Marie actually had a key, we cannot distinguish between cases 5 and 7. If the facts are known but it remains uncertain whether a general legal concept applies to them, there is a gap of recognition, which Alchourrón and Bulygin connect to Hart’s penumbra.189 For example, it may be uncertain whether a delay of two to three minutes for Marie is ‘substantial’ in the sense relevant to Article 701. Finally, the system may provide an answer while treating as irrelevant a circumstance that ought to affect it, which Alchourrón and Bulygin call an axiological gap. For example, say we are dealing with case 5, but we also have evidence that Jean is keeping the gate locked primarily to obstruct Marie. The analysis would still produce the ‘facultative’ solution, even though an actual French court may apply abus de droitabuse of right , resulting instead in a prohibited solution. Similarly, equity or good faith may provide the evaluative reason why a circumstance ignored by the selected rules ought to change the result. It is worth noting that a harsh or unjust answer is not in-of-itself an axiological gap. If the system has recognised every relevant distinction for the given context, but still gives a harsh solution, then no gap exists. It is also worth noting, as they are sometimes confused, that because an axiological gap presupposes an answer exists, the same case cannot also be a normative gap within the same context.190

Read architecturally, Alchourrón and Bulygin give the formalist approaches surveyed in §3 a place within a broader legal order. A normative system in their sense need not be the legal order taken as a whole. For a particular analysis, the applicable norms form its basis, while the properties, cases and solutions fix the terms relative to which the normative consequences and completeness are assessed. This allows an AGIL to place a formal calculus inside a broader concept of law without mistaking one for the other. A positivist account such as Hart’s explains how sources and institutional acts acquire legal authority within the wider legal system. Alchourrón and Bulygin provide the vehicle in which a particular analysis can be derived from those sources. The formalisms surveyed in §3 can then perform different operations inside that structure. Alchourrón and Bulygin do not supply those formalisms. Their contribution is to structure the inputs and outputs of the formal analysis, and the boundaries by which any claim of completeness is made. To give an example, the Kanger-Lindahl calculus can specify one part of the Universe of Solutions rather than a complete normative system. It can generate the normative positions legal agents may occupy, but not which position the law assigns in a particular case. Building on Alchourrón and Bulygin’s conception of a normative system as a deductive relation between cases and solutions, the Lindahl-Odelstad approach gives it a more explicit form by treating the conditions describing a case as grounds, normative positions as consequences, and norms as the glue between them.191 The Kanger–Lindahl calculus can therefore structure the solutions available to the system, while the Lindahl-Odelstad approach can provide the connection by which cases produce those solutions. Normative Systems is valuable to an AGIL not because it displaces broader jurisprudence, but because it provides a way to situate formal approaches in a contained vehicle traversing the broader legal system.

The goal of the work in question is determinative here, as I think it is important to understand all of the works we’ve been describing from the perspective of what they are trying to achieve. For example, I don’t agree with Alchourrón and Bulygin that positivist and realist jurisprudential philosophies are based on an Aristotelian concept of a system, and are therefore, as they seem to imply, somehow out-of-date in light of the scientific revolution’s rejection of Aristotle’s holistic approach in favour of the more ‘modern’ split between empiricism (eg experimental science) and rationalism (eg mathematics). Indeed, in pursuit of their specific goals, they have a tendency to brush past significant methodological questions with claims such as:

Kant’s attempt to reconcile rational science with empirical science and thereby restore Aristotle’s unitarian ideal was not succesful [sic], so that the clear-cut division between rational and empirical sciences persisted until the end of the nineteenth century.192

What this account leaves out is the extent to which the ‘modern synthesis’ they go on to describe retained a distinct Kantian methodological inheritance. Albert Einstein, perhaps the signal figure in the modern synthesis, brought rational and empirical considerations together in an approach influenced by Kant, even if it differed in its conclusions.193 I appreciate such broad brushes are debatable. Indeed, I’m sure this essay is vulnerable to similar ripostes at various points, but that is indeed the point. In any event, I would turn Alchourrón and Bulygin’s own logic back on them. Their conception of a normative system is constructed for their particular work in legal logic. The other works of theory about legal system(s) that I’ve mentioned in this essay are engaged in philosophical or sociological projects somewhat separate from theirs, indeed, often explicitly to describe a legal order holistically. Casting that holistic approach as outdated in light of advancements in science or logic as they do is really to make a series of unargued-for claims about philosophy and language. Consider, for instance, this observation of Bulygin in 2015:

In current legal theory the concept of a system is frequently understood as being necessarily complete (Kelsen, Zittelmann, Cossio) or else necessarily incomplete (Kantorowicz, realist movements). This is a consequence of not taking into account the relational character of the concept of completeness, and of regarding legal orders as wholes, embracing all legal norms, instead of considering each normative system separately.194

To contextualise what Bulygin is claiming here we need to dwell on Alchourrón and Bulygin’s concept of completeness a little longer, particularly as it relates to their discussion of interpretation. Completeness, in their accounting, can be assessed only after interpretation has fixed the norms that constitute the system. They explain that a change of interpretation which alters normative consequences can be represented in either of two ways. One treats it as replacing a norm in the system’s basis, meaning the set of interpreted norms from which its consequences are derived. The other treats it as changing the rules used to derive those consequences. Either representation yields a different normative system because its consequences have changed. For their own analysis they only want to deal with logical rules of inference. They therefore use the first representation and treat a change of interpretation as the replacement of one interpreted norm by another.195 Alchourrón later rearticulated the first kind of interpretation by distinguishing the legal texts that form the ‘Master Book’ from the norms attributed to those texts through interpretation, which constitute the ‘Master System’. The same texts may therefore generate different normative systems, and completeness can be tested only after one has been selected.196 Where a term is penumbral, the interpreter must decide whether the individual case falls within the interpreted category. Deduction alone cannot tell us whether, for the purposes of the rule, a bicycle should be treated more like a motor car or more like a pair of roller skates. Where an interpreter identifies an axiological gap, the question is whether a property ignored by the existing rule ought to affect its answer. An interpreter may support either judgment by drawing an analogy, following precedent, invoking a moral principle, or appealing to statutory purpose or legislative intention.197 If the interpreter responds to an axiological gap by treating the rule as defeasible, the operation is to introduce an unstated exception and replace the broader rule with a narrower one. The revised rule can then support ordinary deduction in the case at hand, although it remains open to later revision.198 Formalisation can expose the point at which the choice enters and trace its consequences, but it cannot make the choice disappear.

An example may help to illustrate how this discussion of gaps, completeness and interpretation can play a role shaping how a computational law project handles law’s brambles. In §3 I briefly mentioned two recent studies that used LLMs in a case-factor analysis. Both studies concerned cases that turned on whether facts arising during a traffic stop gave an officer reasonable suspicion of drug trafficking sufficient to detain a motorist. In the first study, researchers had LLMs read those cases and recover candidate factors bearing on their outcomes.199 In the companion study, they worked from an existing expert-defined factor representation, and used LLMs to identify magnitudes for the factors where degree mattered and then used those magnitudes in a rule-based construction of legal arguments.200 Such a set of factors can be treated as a candidate Universe of Properties in Alchourrón and Bulygin’s terms, the features by which cases are sorted and outcomes assigned. Recovering the factors that bear on a run of decisions is an attempt to reconstruct what Alchourrón and Bulygin call the thesis of relevance, the properties the law has been treating as relevant.201 One of the factors from the traffic stop cases is ‘Legal Indication of Drug Use’ which is determined by whether the defendant has a past drug conviction. The factor is described as binary, insofar as severity of the prior crime is not relevant. Nevertheless, it is plausible that a new case may arise where a lawyer may need to argue that a single minor drug conviction from twenty years ago is not sufficient to support ‘Legal Indication of Drug Use’ as grounds for detainment. If the recency of a prior conviction mattered, it would introduce a new property and potentially reveal an axiological gap, a case the rules cover on their face but which an excluded property ought to govern. In Alchourrón and Bulygin’s terms, this is a difference between the thesis of relevance and a hypothesis of relevance.202 A system just using the recovered factors can associate a case with others where a prior factor was present, but it is unlikely to raise the distinction a lawyer might, that a factor no case has considered relevant ought to decide a new one. Making a legal argument or advising a client about possible outcomes is sometimes knowing where the law may be open. The risk is not a wrong answer but an unflagged choice. When such an approach is used in prediction or argument, the more it fits the cases already decided, the more complete it may appear. But fitting how cases were settled is not the same as establishing what ought to matter in a new case.

Lindahl and Odelstad’s theory of joining-systems gives us an additional way of thinking about where any apparent completeness breaks down in this particular example.203 They distinguish between a stratum of grounds, a stratum of consequences and the joining relation between them. They allow a legal concept to mediate between the grounds and the consequences as an intervenient. ‘Reasonable suspicion’ can be represented as an intervenient between the facts of the traffic stop encounter and the officer’s normative position, including whether the officer may detain the driver. The factor studies work mainly on the grounds side of the Lindahl-Odelstad approach. We can use such a factor-discovery method to recover from decided cases that a past drug conviction bears on reasonable suspicion. As noted, in the expert-defined representation used by the second study, that factor is treated as binary, while categorical magnitudes are added only to the four factors for which degree appears to matter. Yet, as discussed, the represented grounds may not necessarily be exhaustive. Whether a conviction from twenty years ago should count in the same way as a recent conviction is a plausibly open question about the conditions under which reasonable suspicion — the intervenient — arises, even if no case in the dataset has yet drawn that distinction. On the consequence side, the studies fix the principal outcome as whether further detention was justified, while a more general analytical approach may need to distinguish the scope and duration of what the officer was permitted to do. In other words there is a connection between the grounds for the stop and the actions of the officer via the intervenient of reasonable suspicion. An AGIL could retain the candidate factors and, where applicable, their categorical magnitudes, while recording that the analysis is complete only relative to the grounds and consequences represented in the selected cases. It could distinguish an omitted or contested ground from an omitted or contested legal consequence. Used in this way, the formalism cannot decide whether recency of a prior conviction ought to matter or which interpretation should prevail, but it should prevent that choice from disappearing inside a prediction or generated argument.

None of this makes those two case-factor studies inherently flawed or without utility. They are carefully narrow, held to a single domain and fixed sets of cases. They are also clear about their work’s limitations, the first study noting the weaker performance of the fully automated pipeline and the need for human involvement and legal knowledge, and the second confining its results to the cases studied, and subjecting the LLM outputs to a detailed error analysis. At the same time, the researchers’ stated ambitions extend beyond the experiments themselves. The first study purports to offer a more efficient alternative to the costly manual identification of factors, and the second seeks to help legal professionals interpret cases and construct arguments resembling those made by lawyers and judges. Indeed, it is because the studies’ techniques are potentially useful in those ways that an AGIL must ask how the approach will generalise and what they may leave out. A narrow scope is methodologically useful because it controls the number of variables and makes evaluation possible, but it also brackets relationships with other concepts, doctrines, institutions, actors and goals that must be brought back when the technique is applied in practice. Similarly to computational law projects in the 1980s, today’s Law and AI projects can sometimes operationalise a conception of legal relevance and reasoning without articulating the concept of a law or a legal system in which it belongs.204 That is more understandable if the work stays purely in the experimental register. But as soon as it leaves that register, it needs that conceptual architecture, particularly if it’s going to be used by other agents, human or computer, with minimal context about its limitations. A key part of any legal training is creating a conception of the role of the agent, whether human or computational, and a certain concept of the legal system in which they are operating. If computational systems are to take on more of a lawyer’s work without those conceptions, neither the system’s narrow focus, nor the fact that it’s hooked up to a powerful LLM, will save it from flawed legal analysis, particularly when that analysis is applied at scale, and particularly when applied to areas of law with more gaps in the senses we’ve been discussing. Even if a lawyer spends their entire career focused on a narrow area of a legal field with few gaps, their understanding of the legal system more broadly and their role in it is critical to the execution of their role, as it informs the choices they make and the way they present their arguments and advice.

Returning to the question of the ‘completeness’ of a legal system, it is here that Alchourrón and Bulygin’s contextualisation of formalist approaches within bounded normative systems meets its conceptual boundary with an account such as HLA Hart’s. Alchourrón, Bulygin and Hart are all analytical positivists, but with different focuses. Alchourrón and Bulygin approach the legal system through the architecture of logic, while Hart approaches it through the practices by which law is recognised and applied. Normative Systems is part of the ‘Library of Exact Philosophy’, whose stated aim was to keep alive ‘the spirit, if not the letter’ of the Vienna Circle.205 Needless to say there is a strong insistence on semantic hygiene running through Normative Systems which manifests in their discussion of completeness:

Some writers do in fact speak in this context about the incompleteness of law. This is equally misleading since the term ‘incompleteness’ suggests the absence or lack of something. But cases of penumbra do not arise because the law is lacking in something: if the system is complete in the sense that is [sic] solves all the cases in the [Universe of Cases], then it also solves all individual cases, which does not exclude the possibility of the appearance of cases of penumbra. But the latter do not arise from an insufficiency or fault in the normative system; they are due to certain semantic properties of the language in general.206

Through their footnotes it is clear that Alchourrón and Bulygin are thinking of Hart, particularly his brief discussion of analytical positivism and completeness in the context of a broader discussion about the distinction between law and morality.207 Indeed, I would note here that that distinction, between law and morality, is really what Hart has in mind when he uses the term ‘complete’, which can be further seen in his later writing in response to Ronald Dworkin:

the law in [cases of penumbral language] is fundamentally incomplete: it provides no answer to the questions at issue in such cases. They are legally unregulated and in order to reach a decision in such cases the courts must exercise the restricted law-making function which I call ‘discretion’.208

I would submit that, perhaps somewhat ironically, what strikes Alchourrón and Bulygin as a tussle about the use of the term ‘completeness’ arises from its penumbral quality. It is perhaps doubly ironic insofar as Hart took his concept of ‘open texture’ from Friedrich Waismann, a member of the Vienna Circle, as I noted above. Circles within Circles. In any event, in Hart’s case, he is claiming that the law is incomplete primarily in contradistinction to claims such as Dworkin’s that there are right answers to legal questions. One might say that Hart is saying that the pre-existing law is incomplete, insofar as a judge applying a penumbral rule has to make a real choice, thereby making new law. Dworkin rarely uses the term ‘complete’, nevertheless we might infer some concept of holistic completeness for the law from the ‘right answer’ thesis and how he applies it with his example of the judicial Hercules. In any event, to the extent that ‘completeness’ plays a role in the Hart-Dworkin debate it is as a contested descriptive term about the limits of law as it presents itself to lawyers and judges, particularly in hard cases. In Alchourrón and Bulygin’s case they seem to be arguing that a legal system qua normative system can be, and ideally should be, logically complete.209 For Alchourrón and Bulygin, logical completeness is:

quite independent of political, moral or philosophical considerations. It is a purely rational ideal, being closely linked with explaining, with the giving of grounds or reasons, which is the rational activity par excellence.210

Untangling the difference between just these three different senses of the term ‘complete’ would take more space than I have here, if it is even possible, or worth doing. Indeed, this tussle brings to mind Wittgenstein’s discussion of the completeness of language in the Philosophical Investigations:

[A]sk yourself whether our language is complete;—whether it was so before the symbolism of chemistry and the notation of the infinitesimal calculus were incorporated in it; for these are, so to speak, suburbs of our language. (And how many houses or streets does it take before a town begins to be a town?) Our language can be seen as an ancient city: a maze of little streets and squares, of old and new houses, and of houses with additions from various periods; and this surrounded by a multitude of new boroughs with straight regular streets and uniform houses.211

For present purposes I would submit that, despite tussles like this one over the meaning of ‘completeness’, Alchourrón and Bulygin’s analytical reconstruction of bounded normative systems can sit within a broader positivist computational system that grounds the existence and content of law in social facts.

Sometimes the reason debates between different positivists, formalists or realists seem interminable or thorny is because the interlocutors are often speaking from the perspective of different members of the same family, or from different suburbs of the same town, with similar but slightly different background assumptions about what they’re debating. Other times, the different perspectives do indeed lead to substantively incompatible outcomes. Hart gestured at this kind of more substantive mismatch in the postscript to The Concept of Law:

Legal theory conceived in this manner as both descriptive and general is a radically different enterprise from Dworkin’s conception of legal theory (or ‘jurisprudence’ as he often terms it) as in part evaluative and justificatory and as ‘addressed to a particular legal culture’, which is usually the theorist’s own and in Dworkin’s case is that of Anglo-American law.212

In the ‘debates’ between Hart, Dworkin, Alchourrón and Bulygin, Llewellyn, or any of the other figures and works from the last few hundred years of legal theory, the question for an AGIL is not their theoretical conflicts per se but conflicts between the substantive outcomes that follow from their theories for the AGIL qua operational system. Some people don’t work as members of the same family, and some suburbs don’t work as part of the same town. An AGIL may draw on more than one of these traditions, but only if it states the purpose for which each is being used and makes any resulting tensions visible. The question for an AGIL is not which theory is most persuasive in the abstract, but which concepts retain a functional utility as a part of a coherent system.

4.3 Navigating the mess

To the extent that there are tensions between logical formalism and a social account of law, an AGIL’s core needs to be both sufficiently transparent such that the tensions at its edges can be clearly seen, and sufficiently flexible such that both the core and its boundaries can be adapted to the type of user and the purpose of the analysis. The relationship between the core, the system, the user and the analysis is not inherently relative, but it is inherently pluralistic. A law student, a solicitor, an academic, a regulator, and a judge do not all look at laws, the legal system, or legal outcomes in the same way, and do not engage in legal reasoning or legal judgment for the same ends. A solicitor advising a client typically wants the prudent course and to know the risks and opportunities that come with it. A regulator drafting a rule wants it to hold across cases not yet arisen. A judge wants to decide the case and justify the decision to those it binds. Joseph Raz locates this plurality in the standpoint of a legal statement rather than in the law itself. He noted that Hart's dichotomy between internal and external statements tends to obscure a third category, one lawyers often use.213 A legal statement may be committed, endorsing the legal point of view as the speaker's own, as judges typically do. It may be detached, made from that point of view without committing the speaker to it, as when a Catholic expert in rabbinical law tells his observant Jewish friend what to do. Or it may be external, a report or prediction about the legal system and its possible outcomes, such as a lawyer’s advice to a client. These points of view are not exclusive to individuals or roles. A lawyer may adopt a number of them in a single sitting. Llewellyn's ‘law-jobs’ suggest a further teleological element, recasting law as a set of recurrent problems that arise in any group, from the family to the state, disposing of trouble-cases, channelling conduct in advance, allocating authority, directing the group as a whole, and juristic method.214 An AGIL can keep its account of the law answerable to the relevant sources while adapting the form and practical bearing of its analysis to the user and task. What varies is not the law being stated, but the job the statement is asked to do. This interleaves with Alchourrón and Bulygin’s account of reconstruction, in which differently configured sets of norms may constitute alternative formulations of the same normative system when they yield the same solutions for the same cases, allowing an AGIL to adapt its formal representation to the task without changing the solutions it yields for the cases in view.

It follows from this that an AGIL cannot assume that there is a right answer to a legal question. Rather, it must assume that there may be multiple right answers that are contingent on the context of the question, the actors involved, the nature of the questioner, and when the question is posed. This is not to definitively comment on jurisprudential correctness of Dworkin’s right answer thesis, it is to note that, for the purposes of the project of AGIL, the right answer thesis cannot underlie its understanding of a legal system, unless it confines itself to a single type of user, in a single context, at a particular time. Dworkin began his book Law’s Empire with an elegant invocation of the spirit I gestured at in §1:

We live in and by the law. It makes us what we are: citizens and employees and doctors and spouses and people who own things. It is sword, shield, and menace: we insist on our wage, or refuse to pay our rent, or are forced to forfeit penalties, or are closed up in jail, all in the name of what our abstract and ethereal sovereign, the law, has decreed. And we argue about what it has decreed, even when the books that are supposed to record its commands and directions are silent; we act then as if law had muttered its doom, too low to be heard distinctly. We are subjects of law’s empire, liegemen to its methods and ideals, bound in spirit while we debate what we must therefore do.215

One of the primary moves Dworkin makes to position and develop his distinct approach to this empire is to posit a judicial ‘Hercules’, ‘a lawyer of superhuman skill, learning, patience and acumen’, who could find the right answer to any problem in law’s dominion, even its hard and complex cases.216 In order to exercise his superhuman judgment in the midst of the variegated legal territory we’ve sketched, Hercules must not only know the map, he must master the territory in such a fashion that in any given case he can synthesise law, fact, reason and morality with clear, coherent and consistent integrity.

There is an interesting analogy between Dworkin’s Hercules and the role some may imagine for some future, greatly enhanced, LLM as a standalone general legal AI without the kind of symbolic framework I’ve been suggesting is needed (herein the ‘LLM-alone’ approach). Before using Dworkin’s Hercules as a vehicle to think about the LLM-alone approach, I’d first note a few similarities and differences between them to ground the analogy. Firstly, both Hercules and LLMs produce outputs in the form of inferences, and deliver ‘answers’ based on inputs, eg the questions raised by a case. No matter how powerful LLMs become, they will still be used via inference, and their utility will still be relative to the ‘rightness’ of their answers. Secondly, an LLM’s distinguishing feature is its generality. It is the product of a vast and general body of material, and it is that breadth, not any bounded or specialist training, that produces its value. Hercules is generalist in a structurally similar way, in that it is his use of a broad and general dataset that distinguishes Dworkin’s account from its positivist rivals. Where positivism confines the judge to the artificial reason of the law, Hercules reasons from the whole ‘great network of political structures and decisions of his community’, and from the political morality that would best justify it.217 His power, like an LLM’s, lies in the breadth of what he draws upon and in his capacity to synthesise that general material into specific applications of legal judgment. There is a difference between them in how they use this general dataset in that Hercules reasons with ‘integrity’, testing each part against a justification of the whole, whereas an LLM distributes attention across the tokens of its context window in a forward pass over billions of parameters set in training by gradient descent. Moreover, it is difficult to say with any confidence how any particular inference is going to perform that operation, whether it will be a simulacrum of integrity, conventionalism, pragmatism, or something else. Nevertheless, the generality of the grounds of their reasoning is what distinguishes them from their respective competitors. Thirdly, Dworkin’s account turns on the distinction between the reasons that genuinely ground a decision and the reasons a judge in fact states for it, with Hercules being the limiting figure in whom the two coincide, his stated justification being the theory that decides the case. As we noted at the start of §4, current LLMs effectively sit at the opposite end of this spectrum as what they actually take into account is fixed neither by their prompt nor by the reasons they give in their output. Let us grant, however, for the sake of argument, that some future LLM overcomes this, reasoning in the way it is asked to and the way it says it does, so that its stated reasons can be trusted as its operative ones. Dworkin works through what this approach to law would entail, discursively, across Law’s Empire, but to take one representative example:

Law as integrity, then, requires a judge to test his interpretation of any part of the great network of political structures and decisions of his community by asking whether it could form part of a coherent theory justifying the network as a whole. No actual judge could compose anything approaching a full interpretation of all of his community’s law at once. That is why we are imagining a Herculean judge of superhuman talents and endless time.218

Taken as a whole, Law’s Empire suggests that the judicial Hercules needs more than mere intelligence to master the map and the territory of law. Hercules needs to exist outside of time and space, so that a single intelligence can apprehend how the whole system, which effectively amounts to most human activity within a jurisdiction, synthesises into the ‘right’ answer. This is Hercules the demigod, who has mastered his labours and ascended into the stars. I read Dworkin’s use of Hercules in the same vein as Rawls’s use of the original position or Descartes’s use of his demon, not as a claim that Hercules is empirically feasible, but as what Kant called a regulative idea of reason, namely a focus toward which the practice of law strains and which it presupposes, but never treats as empirically constitutive of the practice or practitioner.219 By considering the ideal, one may illuminate the real. Dworkin did not frame Hercules in Kantian terms, but the framing matches what seems to be his intent. However, when read as an analogy with an LLM-alone general legal AI, Hercules is recast from a regulative idea to what seems to be empirically necessary in some sense if such an approach were to be realised. In that light he reads more as a kind of reductio ad absurdum against the notion that AGIL could be achieved by an appropriately trained and refined LLM in isolation, reasoning to right legal answers. This is not a comment on the technical feasibility of gathering sufficient data, or of devising training and refinement good enough to produce high-quality answers to particular legal questions. It is a reductio against the sufficiency of the LLM’s structure to fill the role the LLM-alone approach to general legal AI would cast it in. The LLM-alone approach asks the LLM to perform the labours of Hercules, to hold the whole system in view and resolve it, while giving it no separable representation of that system an auditor can inspect, contest, or verify an answer against, independently of the inference itself. Even granting the faithful LLM of the future posited above, whatever such a model models about the relevant legal system(s) (however defined) stays entangled in the same process of inference that produces its answer, which cannot be audited. A neuro-symbolic AGIL like the one we’ve been sketching in this essay makes no pretence to Hercules’s all-seeing vantage. It does locally and auditably what Hercules is imagined to do globally and ideally, navigating the territory with a map rather than attempting to see the whole of it at once, as if it were a god.

Let’s bring together the various strands of an AGIL’s approach to the brambles of law, such as equity, good faith, open textured rules and standards, and so-called legal realism. The first requirement is a clear concept of a legal system, since without one the system has no stable map of the relationship between its formal reasoning and the social and institutional practices in which law exists. I would suggest a positivist philosophy around its formalist core, for example Hart and Raz paired with Alchourrón and Bulygin, conscious of the tensions that pairing involves. Whatever choice you make, there must be a clear relationship between whatever formal logic the AGIL employs, such as the non-monotonic deontic logic I mentioned in §3, and the liminal zones between law, morality, politics and society. The primary utility of Alchourrón and Bulygin is their account of completeness and their taxonomy of gaps, which supplies a vocabulary for identifying the limits of the formal logic. The formalist approaches surveyed in §3 can then be encapsulated within such a system rather than mistaken for a complete theory of law. Beyond the formalist core, the work frameworks like those of Hart and Raz’s are playing is to stake out a cogent jurisprudential position for the AGIL vis-à-vis other approaches such as realism, natural law or interpretivism, particularly as it seeks to serve different types of users. It is hard to overstate the importance of such choices in shaping the technical choices you make in logic, architecture and the role of LLMs, as seen in the example of how you draw a set of properties from a line of cases which you then use to generate advice or arguments. That is just one small example, which plays out many times in many different ways throughout the system. Once you have an idea of the AGIL’s concept of a legal system you must situate it within the contexts and uses to which it will be put, such as the production of advice, comment, assessment, or even judgment, and the various actors who may be using it, including lawyers, regulators, laymen and perhaps even judges. Legal statements are not fundamentally relative. However, they are pluralistic in their interpretation, and an AGIL must account for that pluralism if it is to navigate the territory’s brambles.

5. How to address law’s problem

One of the schools of Tlön goes so far as to negate time: it reasons that the present is indefinite, that the future has no reality other than as a present hope, that the past has no reality other than as a present memory.220

As I discussed in §1, in the context of the vast reach of the modern regulatory state, a great deal of the work that goes on inside a bank, an insurer, a health system, a government department, a legislature, a defence force, a university, or a business is legal reasoning being done under the pressures of time, scale, and resources. Whether a contract can be entered into and on what terms, what follows if the counterparty defaults, whether a new product or policy sits inside the existing framework or outside it, whether the decisions of the last quarter or session have stayed consistent with duties to customers or citizens, shareholders, the regulator, and multiple overlapping regulations, statutes, opinions and guidance. Such questions arise and are acted on, or sometimes neglected, every day, in enormous volume, by in-house counsel, outside counsel, compliance teams, product and risk teams, policy advisers, operations staff, and diverse others, many of whom are not lawyers. The traditional model of legal services, even with the latest AI tools at hand, simply cannot deal with the scale and depth of this reality, as has been evident for some time, as we can see in the various statements, metrics, and case studies to that effect I presented in §1. I suggest that our current AI moment presents a real opportunity to tackle this long-standing problem facing modern legal systems. AI is not just compliance dashboards, chatbots, enhanced research tools, or ‘agents’. It can also be accurate and auditable computational legal reasoning integrated into the systems powering the modern economy and state. When structured correctly, it can operate at the speed and scale of the technologies those diverse actors are already using, or the even more sophisticated technologies and intelligences they will be using in the near future, and is better suited to deal with the cumulative messiness all of that speed, scale and variety entails. I have started to sketch how a neuro-symbolic general legal intelligence grounded in jurisprudential philosophy can achieve this.

5.1 Making computers follow laws?

There have been a number of recent papers in the field of ‘Law and AI’ advancing what initially appears to be a different approach to legal AI, arguing for some version of using law to ‘align’ LLMs and agents (herein ‘legal alignment’). ‘Alignment’ here typically means getting an LLM to behave reliably in ways that match human intentions, however those intentions are defined. An early paper in this vein was John Nay’s ‘Law Informs Code’, which cast law as ‘the applied philosophy of multi-agent alignment’.221 Nay developed a threefold test for assessing the utility of any alignment framework. First, it should be already constituted with ‘modular constructs built to handle the ambiguity and novelty’ alignment involves.222 Second, it should be both useful for today’s AI and scalable to future AI. Third, it should be rigorously tested such that it has abundant data for the alignment of LLMs. Nay goes on to make some subsidiary arguments about alignment per se, and then briefly sketches some case examples in contracts, legal standards, and fiduciary law. In a 2025 paper, Cullen O’Keefe, Ketan Ramakrishnan, Janna Tay and Christoph Winter consider the dynamics of the relationship between humans and AI agents, and the ‘risks to life, liberty, and the rule of law’ AI agents may pose, particularly in high-stakes settings.223 They argue that AI agents should be trained to be loyal, but only within the bounds of the law, refusing illegal acts even where these would serve the principal. They contrast a ‘Law-Following AI’ with an ‘AI henchman’ who treats prohibition as a bare cost, suggesting an analogy with the distinction between acting from Hart’s internal point of view and acting like Holmes’s bad man.224 They argue that language models weaken the premise of earlier formalisation efforts:

Today’s frontier AI systems can already reason about existing natural-language texts, including laws, with some reliability—no translation into computer code required. They can also use search tools to ground their reasoning in external, web-accessible sources of knowledge, such as the evolving corpus of statutes and case law. Thus, the capabilities of existing frontier AI systems strongly suggest that future AI agents will be capable of the core tasks needed to follow natural-language laws, including finding applicable laws, reasoning about them, tracking relevant changes to the law, and even consulting lawyers in hard cases.225

They argue that AI agents should be designed to find, interpret and follow natural-language law, rather than be governed by a small set of specific legal commands translated into formal code.226 This would entail treating agents as ‘legal actors’ capable of bearing duties without necessarily being granted rights or legal personhood, while also imposing duties on their developers and principals.227 Finally, they suggest that ex ante regulation of agents and their developers needs to be taken seriously due to the inherent difficulties in both control and enforcement.228 Noam Kolt’s ‘Governing AI Agents’ can be read as a companion of sorts to ‘Law-Following AI’, analysing AI agents through both the economic theory of principal-agent relations and the common law doctrine of agency.229 Interestingly, Kolt also observes that these lenses fail in significant ways in the realm of enforcement.230 Indeed, reading these two papers one is left with the impression that the utility of the common law of agency is sufficiently circumscribed by the enforcement problem such that pursuing its application to AI agents further is an intellectual cul-de-sac. Jack Boeglin’s recent paper ‘Aligning Artificial Intelligence to the Law’ treats law as a ‘coach’, using concepts from agency, contract and public law to frame a discussion around control, whose interests it serves, how it interprets instructions and what limits bind it.231 One of Boeglin’s observations is particularly pertinent to AGIL:

For advances in computational law to best further AI alignment, a shift in focus is likely necessary. Current benchmarks for assessing an AI system’s legal understanding often look to whether it can correctly identify the legal consequences of externally described circumstances (e.g., if X does Y, does that violate Z law?). But legal alignment requires more than the ability to answer legal hypotheticals; it calls for an assessment of AI’s own circumstances. This might require AI first to learn to reliably describe its own conduct and place in the world from an external perspective—a task that might require a greater degree of self-awareness than AI systems currently appear to possess. Legal alignment thus throws down a gauntlet to computational law, as well as offers yet another reason to think that the field holds great promise: it could help to solve the alignment problem.232

I agree computational law can help to solve the alignment problem, and have tried to reach down to pick up the gauntlet in this essay, albeit I would note in passing that I disagree with the suggestion that legal alignment might require more self-aware AI to start to make sense. I would also briefly note that in these various papers, ‘law’ is often not defined as such and the reader is often left to fill in a few gaps as to what the author really means when they refer broadly to ‘contract’ or ‘agency’ law. Reading them from the perspective of a non-American legal background, one is reminded of HLA Hart’s aside about the unstated cultural limits of Dworkin’s work (quoted above), ie that the theories are often addressed to a particular legal culture, usually the theorist’s own. Unstated cultural scopes become more relevant in an inherently borderless technology like an AI system.

Perhaps the most prominent recent statement of the legal alignment agenda is ‘Legal Alignment for Safe and Ethical AI’ (the LASEAI paper), which argues that AI alignment has overlooked law as a resource.233 The authors note that current alignment practice typically relies on opaque, and often conflicting, private artefacts.234 Against these they make similar arguments to those of prior papers, noting that law is a more democratically legitimate, concrete and tested framework. They set out three research directions: using laws as normative content for alignment, using legal theory as a guide for AI decision making and using legal principles such as fiduciary relationships as a ‘blueprint’ for an alignment framework. They then identify several open questions. First, defining what ‘law’ means as a target for alignment. Second, determining how legal alignment will work considering laws are not written for that purpose, how to actually make LLMs follow laws, and considering the circularity risk if AI becomes involved in making laws. Third, considering whether legal alignment will suppress innovation, whether it could be gamed, and whether it will scale. Like O’Keefe and colleagues they claim that legal alignment has become possible because:

Advances in language modeling have dramatically improved the legal capabilities of AI systems. Unlike prior efforts to computerize law that relied on the formalization of legal rules … language models have enabled AI systems to reason about law in the natural language in which law is constituted and communicated.235

I would pause here to note that one of the inherent challenges of both legal alignment, and Law and AI more broadly, is the multi-disciplinary scope of the undertaking. Reading papers such as LASEAI can at times feel like reading the contents page to a multi-volume work that doesn’t exist, as if one had slipped into a Borges story. Perhaps one might feel similarly reading this essay. There is a fine line between setting out a broad agenda, or making a broad argument, and leaving so many fundamental questions unanswered that the agenda or argument becomes a collection of unobjectionable claims, stretched so thin that they dissolve on contact.

In any event, it is important to try to ground the discussion around legal alignment in some realities. Firstly, I’d note that an ‘AI agent’ is a neuro-symbolic system that wraps an LLM in tools, validators, planners, stores and memory, and other symbolic structures. An AI agent is typically a flexible and versatile software program, but, outside of the LLM itself, it is a regular program with symbolic rules. Secondly, in my experience working as a software engineer within technology enterprises, ‘agentic AI’ in such contexts is largely about making specific parts of what was already a highly structured workflow more efficient and scalable. The areas of software development in which agents have the greatest utility are those which have clearly defined inputs and outputs and a high level of structure, such as the development of frequently produced modules with clearly defined patterns such as user authentication or dashboard interfaces, or the monitoring and remediation of tests within a continuous integration environment. It seems likely that AI agents will continue to progress in capability such that, within such sufficiently structured workflows, they achieve capabilities comparable to or beyond what a human engineer can achieve, at least in specific parts of those workflows. Nevertheless, it also seems likely that, however capable LLMs become, the quality and nature of the structures in which they are embedded — whether those structures are called ‘agents’ or something else — and the use to which they are put within those structures, will be key to both their utility and integrity in that context. I am not foreclosing the possibility that LLMs become better at reasoning in expert domains such as law, merely observing that their deployment in operational workflows in the area in which they have the greatest expert capacity, namely software engineering, suggests that the quality of their output is fundamentally tied to the quality of their operational structures, in particular the ability of those structures to model the existing workflows of the technical expertise they seek to emulate. Indeed, it was not reading the literature of formal logic in law that initially suggested to me a neuro-symbolic approach was the path toward AGIL. It was a decade of hands-on experience in software engineering, the field in which AI agents are undoubtedly the most capable, and have already revolutionised the day-to-day of the industry. If future LLMs become as proficient at legal reasoning as they currently are at software development it would be a significant step change from their current capabilities, and yet few professional software engineers are presuming that AI agents should be left to code by themselves on professional projects without sophisticated symbolic harnesses, and it still seems possible they never will. In any event, once we ground what an agent, or agentic workflow, is in practice, the distinction between ‘agentic AI’ and ‘symbolic AI’ fades away vis-à-vis their current use and application in a professional context. To put it another way, ‘agents’ and the structures they are effective in are symbolic structures. The current question is not whether LLMs deployed in a legal context will use an external symbolic structure to guide their behaviour, but what external symbolic structures are used to guide their behaviour. There have never been, and may never be, free-floating LLMs reasoning to right legal answers by themselves like some Dworkinian Hercules, whether as Hercules the hero who conquers his labours, or Hercules the villain who goes mad and kills his wife and children. The discussion of legal alignment should try to stay tethered to the reality of what LLMs and ‘agents’ actually are, and how they’re actually used.

5.2 How to talk about the future

More broadly, it is helpful to consider the structure of some of the arguments underlying the legal alignment papers. Firstly, we should consider what Alfred Nordmann called, in the context of the debates around nanotechnology in the 2000s, the ‘if-and-then’ argument.236 In short, the legal alignment version of this argument is: let’s assume some future LLMs or AI agents will be significantly better than the ones we have now, and then let’s consider what this means about the relationship between law and AI. Nordmann argued that arguments like this often start with a speculative if, and then the conditional quietly hardens into a when, such that a possible future is treated as a settled premise from which present obligations can be deduced. The problem, Nordmann argued, is not that speculation about the future is illegitimate, but that this particular structure lets the hypothetical do the work of the actual. He called the move a ‘radical foreshortening of the conditional’,237 in which ethics, or in this case law, is made to leap ahead of the science, fastening on vivid scenarios that may never arrive while the real and pending questions go unattended. The consequence is that ‘the imagined future overwhelms the present’, redirecting scarce analytical attention toward imagined implications and away from the systems in front of us.238 As well as making similar points about the debate around nanotechnology,239 Armin Grunwald adapted the language of Niklas Luhmann to frame the problem as a kind of Zen koan: that we never have access to future presents, only to present futures.240 Yana Suchikova has applied this thinking to our current AI moment, arguing that if-and-then arguments structure much of the debate and draws attention away from the real and urgent challenges already affecting society in favour of distant speculation.241 Within sociology it has been posited that expectations about technological futures are not neutral forecasts but are performative.242 Such expectations cannot be verified in advance because checking such a claim is not really distinguishable from trying to build the actual technology the claim relies on.243 As such, some recommend a shift from looking into the future to looking at it, treating a projected future as an object of critique rather than a place from which premises may be drawn.244 On this view the very indefiniteness of the language of ‘progress’ lets it hover outside of agency and action while still mobilising and legitimating present effort.

To balance these points we might note that anticipating the future is hardly foreign to law or public policy. Law and lawyers often work in anticipation of the future, and would be seen to have failed if they did not. Environmental law is a good example, with the precautionary principle requiring action against serious or irreversible harm before the science is settled, and processes such as impact assessment requiring stakeholders to reckon a project’s likely consequences.245 The same orientation to the future governs how statutes are written, and how they are interpreted. Legislators draft prospectively and, ideally, reach for technology-neutral language that fixes on the function a technology performs rather than the form it happens to take, ideally reaching a technology’s effects rather than its artefacts, holding the contest between rival techniques undistorted, and outlasting the state of the art.246 The industries involved in the production of policy sometimes use a vocabulary of ‘anticipatory governance’, which David Guston described in terms of capacity building, built less from forecasting than from the exploration of multiple futures, public engagement, and sociological work, so that anticipation becomes a rehearsed readiness.247 That preference for readiness over prediction can be seen to be responding to a difficulty that David Collingridge called the dilemma of control.248 The dilemma is that the knowledge needed to control a technology and the power to change it are seldom available simultaneously. In its infancy a technology can be more easily changed or even abandoned, but its eventual social consequences cannot yet be predicted with enough confidence to justify such interventions. By the time any negative consequences are apparent, society and the surrounding economy have adjusted around the technology to a degree that makes any major change disruptive, costly and slow. Collingridge argued for a focus on preserving the ability to change the technology once it has matured.249 A comparable shift from foresight to adaptive capacity can be seen in Lyria Bennett Moses’s account of law and technology.250 She argues that the familiar image of law forever losing a race against technology is misleading, since most technologies fit comfortably within existing legal frameworks and the ‘race’ metaphor can provoke urgent, poorly conceived responses. She contends that future-proofing through drafting alone is insufficient, and that what is needed is a standing institutional capacity to adapt once the future has arrived. What that requires is the regulatory machinery to notice when a rule has grown uncertain in its reach, ill-fitted to its own purpose, or simply obsolete, and to revise it in good time.251

A current in contemporary ethics tends more in the direction of the precautionary principle, holding that where a harm would be irreversible or catastrophic we are obliged to treat its bare possibility as though it were a fact and to act in advance. If a technology could bring about human extinction or permanently foreclose humanity’s future, then even a small probability of that outcome, set against a loss that is effectively limitless, comes to dominate any cost-benefit calculus. Nick Bostrom defines an existential risk as one that threatens ‘the premature extinction of Earth-originating intelligent life or the permanent and drastic destruction of its potential for desirable future development’.252 Hans Jonas had earlier articulated a version of this underlying ethic in the 1980s, arguing that an age of world-altering technology calls for what he named (non-pejoratively) a ‘heuristics of fear’, which consults our fears before our hopes because we come to know what we do not want sooner than what we want.253 Applied to artificial intelligence, the fear crystallises in what Bostrom, similarly to Collingridge, calls the ‘control problem’. Bostrom articulates it as the difficulty of ensuring that a system more capable than its creators remains under their direction and open to correction.254 Stuart Russell argues that this difficulty is inherent in the so-called ‘standard model’ of the field, in which a machine is built to pursue a fixed objective that is specified in advance and handed to it from outside.255 In the standard model any objective simple enough to be written down is likely to leave out something important, and a capable enough system will then pursue the letter of its instruction to lengths its designers never intended. The potential destruction such monomaniacal systems could wreak seems to militate in favour of strong preventative action.

Within the AI alignment community itself, some have voiced doubts about the ‘speculative’ register of the heuristics of fear. In ‘Concrete Problems in AI Safety’, a group of researchers that included two of Anthropic’s co-founders argued that one ‘need not invoke these extreme scenarios to productively discuss accidents’, that doing so ‘can lead to unnecessarily speculative discussions that lack precision’, and that it is ‘usually most productive to frame accident risk in terms of practical (though often quite general) issues with modern ML techniques’ (meaning machine learning techniques).256 They set out five such ‘practical’ problems, from avoiding side effects and reward hacking to safe exploration, robustness under distributional shift, and the scalable supervision of systems whose behaviour is too costly for humans to check directly. Paul Christiano’s idea of prosaic alignment operates on the working assumption that artificial general intelligence may arrive without revealing ‘any fundamentally new ideas about the nature of intelligence’, so that the task is to align the kinds of systems we are already building.257 Practical alignment work includes reinforcement learning from human feedback alongside approaches that use models themselves to help apply written principles, as in Constitutional AI and deliberative alignment.258 The task of supervising systems that come to outstrip their human supervisors remains open, albeit it is being pursued through proposals such as iterated amplification and debate between models adjudicated by a human, as well as empirical work on weak-to-strong generalisation.259 The leading labs have made this ‘empiricist’ stance clear, with Anthropic viewing empirical evidence as ground-truth and the space of possible systems and failures as too large to be traversed ‘from the armchair alone’,260 and OpenAI describing its alignment work as an ‘iterative, empirical approach’.261

In the context of these debates and projections about the future, I suggest that the optimal path for legal alignment is to adapt Kant’s concept of regulative ideas of reason.262 For the purposes of legal alignment it would be more useful to use speculative futures of AI regulatively, as a focus imaginariusimaginary focus ‘directing the understanding … although it is only an idea’, as opposed to constitutively, which is to say assuming its empirical existence and considering what follows from its existence.263 The distinction is subtle, so it’s worth restating: the goal is to use imagined futures as a focus imaginarius futuriimaginary focus of the future to direct the attention toward present projects as opposed to considering what logically follows from a future imagined to be actually constituted. The reason this is a better way of approaching possible futures in legal alignment is because it avoids many of the pitfalls mentioned above, and focuses attention in the right direction without sacrificing the disciplining role of possible futures. It potentially avoids the inherent weakness of legal academics discussing the possible technical trajectories of artificial intelligence in ways unconvincing to a technical audience and misleading to a legal one. It also avoids the ‘if-and-then’ critique and its variations, by not being hostage to a premise being empirically constituted now or in the future. Moreover, it sharpens the focus on the immediate empirical project the regulative idea is being deployed for. To give an example of how this plays out, consider that one rejoinder of the so-called alignment speculators to the so-called alignment empiricists is Nick Bostrom’s treacherous turn, the claim that a capable enough system would behave well precisely while it is too weak to do otherwise, so that good conduct now tells us nothing reassuring.264 Taken as a regulative idea, this can be used as a focus imaginarius futuriimaginary focus of the future in an empirical project to make the inner workings of LLMs more tractable.265 Conversely, if the treacherous turn is treated as a constitutive premise it essentially turns a productive fear into an unproductive paranoia. If you hold it to be empirically the case in some sense that you cannot trust LLMs, then the effort of AI interpretability is inherently a losing game, as whatever you find could be yet another deception. Moreover it raises a host of questions about what ‘treachery’ could possibly mean for billions of parameters tweaked by gradient descent. I suggest the heuristic to use when assessing the proposals of the legal alignment literature is therefore not whether it speculates about the future, but whether its speculation is used regulatively, to engender a present empirical endeavour, or constitutively, as an empirical premise underlying a conclusion or project. With that heuristic in mind, let’s return to the various legal alignment papers I referred to earlier.

Each of the legal alignment papers I mentioned above can be mostly read charitably to fall on the side of using projected futures as regulative ideas as opposed to constitutive premises. However they also sometimes slip onto the other side of the ledger, and at times seem to reason from the premise that an imagined future is empirically constituted. For example, Nay discusses using AI in the role of a fiduciary, which is a recurrent theme in this literature, as follows:

[F]uture more advanced AI is likely to be more analogous to something like a Trustee administering investments in complicated private equity transactions. We should dial up the fiduciary obligations for more advanced AI. Another way of looking at this: assuming increased capabilities, AI could enable fiduciary duties to be more broadly applied across digital services. In scenarios where an agent is trusted to adopt a principal’s objectives, standards that help ensure the agent can be trusted could be foundational to application-specific training processes. In addition to traditionally clear-cut fiduciaries (such as investment advisers), automated personal assistants, programming partners, and other emerging AI-driven services could be designed to exhibit fiduciary obligations toward their human clients. Advancing capabilities of AI could make this possible by enabling scalability of high-quality personalized advice (the basis of the duty of care), while the advancing capabilities make the duty of loyalty component increasingly salient.266

Nay goes on to suggest ways in which an AI system might be trained to exhibit fiduciary behaviour.267 Understood narrowly, the imagined fiduciary can operate regulatively, as a focus imaginarius futuriimaginary focus of the future directing current work on training data, benchmarks and validation, without assuming that advanced AI will emerge, that an AI system will become a fiduciary or that its service will attract fiduciary duties. Nevertheless this elucidates how easily that projection can acquire constitutive force. Nay uses something he contends ‘is likely’ to occur to support something ‘we should’ do. He projects capabilities that make future AI analogous to a trustee, to support the call to ‘dial up’ fiduciary obligations and extend fiduciary approaches across digital services. These moves risk collapsing three distinct propositions: that AI can be trained to exhibit fiduciary behaviour; that an AI-enabled service should be governed by fiduciary standards; and that AI might itself occupy a fiduciary role. These are, respectively, an empirical programme, a normative choice and a contingent future. Keeping them separate preserves the ability for the research to branch in other directions. Research might instead favour retaining a human or institution as fiduciary and using AI only as decision support, imposing duties on developers or deployers alongside auditing and technical constraints, or not automating the function at all. Greater capability might alter that comparison, but it can’t determine it. A regulative idea of the future motivates inquiry while leaving its endpoint open. Presented as one trajectory, the imagined future instead supplies a premise for present obligations while obscuring both the conditional steps and the alternatives. The objection is not to imagining trustee-like AI, but to allowing that image to predetermine the direction and destination of inquiry.

In a similar way, a passage in the Law-Following AI article quoted above moves from the limited present claim that frontier systems can reason about natural-language law ‘with some reliability’ to the forecast that ‘future AI agents will be capable’ of the core tasks required to ‘follow law’, and from there to a practical conclusion about how agents should be designed. The claim that ‘no translation into computer code’ is required can be read narrowly or broadly. If it is read narrowly, as meaning that law need not all be manually translated into code in advance, it is compatible with the AGIL approach I’ve been laying out here. If it is read broadly, as suggesting that formal representations have no place in legal AI, I would suggest that that is jumping the gun in the way I’ve been describing. If we employ such assumptions, neuro-symbolic systems are not explicitly rejected, but they disappear from view, silently narrowing the available design space. As with Nay’s fiduciary trajectory, collapsing the intermediate steps hides the points at which empirical inquiry might favour a different arrangement. Moreover, as I’ve noted above, an ‘agent’ is not simply an LLM but a neuro-symbolic system, and ‘agentic systems’ are often quite substantial symbolic architectures wrapped around those agents. The question is not really whether translation is required tout court, but which legal materials and operations should remain in natural language, which should be formalised, and how the two should relate.

It may seem like I am cherry-picking quotes from the literature in order to make a point, and I am. My main point in this section is one of language, framing and focus. I have chosen specific language and phrasing in which the shift is easiest to see, not because they exhaust the rich and detailed arguments in the papers I’ve mentioned. I am not trying to argue against Law Informs Code, Law-Following AI, or legal alignment more broadly, and I share many of their respective aims. But framing and specificity matter because they change both the focus and the course of the work. Used regulatively, imagined future AI leads us to ask which present project the idea should orient, who should undertake it, and how it should be tested and revised, rather than allowing a predicted future to narrow the approaches considered in advance. For lawyers, regulators and legal theorists, turning law into an alignment framework can therefore be treated as a present empirical programme. The same distinction helps clarify the apparent debate between ex ante and ex post intervention. Possible high-stakes futures can motivate tests and safeguards now without being treated as facts, and regulation is not constitutive simply because it acts in advance. The slippage occurs when projected features of aberrant AI are treated as settled and allowed to harden into particular regulatory or technical conclusions. As Bennett Moses suggests, the better aim is to preserve the capacity to intervene proportionately and revise that intervention as evidence accumulates and systems develop.

5.3 Systems should be the focus

Turning to focus on the empirical project of legal alignment, I will briefly consider the research agendas proposed in some of the literature, taking the LASEAI paper as the touchstone. One could potentially read this current essay as an attempt to answer the first two parts of the first item on that paper’s agenda, namely how to handle law’s brambles and, relatedly, how to handle the distinction between law and morality, in legal alignment. Indeed, I would argue that some version of an investigation like my brief sketches in §§3–4 is required to make legal alignment make any sense. Law is not a science. A ‘research agenda’ with multiple constituent parts and multiple people involved in it that purports to take ‘law’ as a useful discipline in another field is meaningless unless there is a concept of law and a legal system a plurality of those involved agree on. One might respond that failing to explicitly articulate a concept of law and a legal system does not make, say, the analysis of AI agents through the lens of principal-agent relationships meaningless, and that would be fair enough. One cannot always begin by exhaustively defining one’s terms. However, in a project where the stated goal is to take the entire field of ‘law’ as a useful discipline in another field that is inherently borderless, and to also explicitly distinguish law’s utility from other possible alignment disciplines such as ethics, coherently articulating a concept of law and a legal system is, I would suggest, a precondition of the endeavour, without which any other claims or projects are essentially meaningless. The authors of the LASEAI paper note that:

[L]egal alignment is distinct from legal regulation of actors that develop and deploy AI, which focuses primarily on using law to govern the individuals and organizations that produce, disseminate, and use AI systems … By contrast, legal alignment focuses on integrating law and legal methods into the design and operation of AI systems themselves.268

I understand the distinction the authors are trying to draw. However, I would suggest that without articulating some definition of ‘law’, or ‘AI system’ for that matter, legal alignment could equally be seen as how the laws that already bind actors designing and operating AI systems play a role in how they think about that design and operation.269 It has long been the case that most people building software try to build it in such a way that it does not transgress any laws, and the same is true for those training and deploying LLMs and agents. I know that that is not what is being argued for, but that additional distinction is essentially an inference on the readers’ part, not a reasoned delineation in the paper. Unless you give some real substance to the key concepts of ‘law’ and ‘AI system’, the argument for ‘legal alignment’ has no real meaning beyond the pre-existing understanding of the reader. I suggest that the prolegomenon to legal alignment, understood as ‘integrating law and legal methods into the design and operation of AI systems themselves’ has to be a jurisprudentially grounded articulation of a concept of law and a concept of a legal system, and a technically and empirically grounded articulation of ‘AI systems’ qua system, for the purposes of legal alignment. Getting seventeen researchers to put their name to a prolegomenon of that kind may be challenging, but that is the point. Whether the concept of law is a mixture of positivist and formalist or more realist and consequentialist, whether it is American, Anglo-American, Continental, or comparativist, whether it involves legislation, regulations or judgments, essentially where it falls on many of the questions I have considered in this essay, will fundamentally alter what ‘legal alignment’ means. So too will the scope of ‘AI system’, whether it is LLM-alone, agentic, neuro-symbolic, whether ‘system’ includes the sociotechnical architecture the machines are embedded in, and where the limits of it are. Attempting to answer those questions will demonstrate whether and how legal alignment is substantively different from how laws already apply to AI and AI companies, and its use as a discipline in alignment vis-à-vis other options such as ethics. I would argue that legal alignment should adopt similar positions to those I have landed on in this essay, but whether my positions are accepted or not, some set of answers needs to be roughly accepted by a plurality of those involved in at least some portion of the legal alignment project for it to be comprehensible as a discipline to any of its prospective targets, whether they be AI researchers or legal theorists. As noted above, Susskind made an analogous point when reflecting on the initial wave of computational law in the 1980s, and despite their disavowal of those approaches, some Law and AI researchers may be in danger of repeating one of the main weaknesses of that period.

A paper that starts to engage in the kind of approach I’m calling for here is Nicholas Caputo’s ‘Alignment as Jurisprudence’.270 Caputo argues that jurisprudence and alignment face the common problem of governing how decision makers should act in novel situations. To explore this he engages in a substantive comparative discussion of the two, in particular comparing Dworkin’s interpretivism with Constitutional AI and Cass Sunstein’s analogical positivism with the use of case-based reasoning in alignment. There is not sufficient space for me to engage with Caputo’s substantive arguments here, however I would offer some brief observations about the form his argument takes. Caputo does not use future advanced AI as a constituent premise, rather he focuses on a comparison of methodologies from jurisprudence and alignment in order to suggest some methodological improvements for alignment. He also observes that the comparative analysis might work in both directions, given that alignment might offer a proving ground for legal theory.271 I am reminded here of one of the debates about jurisprudence and computational law in the 1980s, which we can see in Susskind:

Professor Bryan Niblett has claimed that “a successful expert system is likely to contribute more to jurisprudence than the other way round.” Our research (see Section VI) and the remainder of this paper casts doubt on that suggestion. In any event, if the majority of the projects mentioned above are indicative of quality, then it is unlikely that many commentaries on expert systems will exhibit the analytical rigour and sophistication of argument that characterise today’s major contributions to legal theory.272

I agree with those writing in the field of legal alignment that law has much to offer AI research, however, besides some exceptions like Caputo’s paper, one sometimes wonders if ‘law’ and legal theory is being over-simplified in order to make an argument more comprehensible to researchers in other disciplines without legal backgrounds.

I would suggest that when both ‘law’ and ‘AI system’ are given substantive definitions, particularly when law is considered in the context of the variety of legal systems and AI is considered in the context of the variety of AI systems, the legal alignment agenda might begin to look more like a complex systems analysis of both law and AI. What I mean by that is that the two may need to be analysed in a similar fashion to how some have thought about applying systems thinking to regulatory governance,273 or in a similar fashion to how Kolt and colleagues recently suggested that complex systems science could be applied to AI governance,274 or in a similar fashion to how Laura Weidinger and her colleagues at DeepMind applied systems thinking to AI safety itself.275 Using systems analysis as the overarching framework for legal alignment has the advantage of employing a language and conceptual framework that is plausibly comprehensible to both legal and AI theorists, insofar as it has already been used by some to separately analyse both disciplines, and is sufficiently capacious to be interoperable with both abstract theory and concrete technical details. I would emphasise that using systems thinking as a meta-framework does not obviate the need to engage in the more in-depth jurisprudential grounding as seen in Caputo’s paper, or in the more in-depth technical grounding one typically sees in non-legal AI alignment papers. Indeed, the point is to give such analyses the space to go in-depth while remaining relevant to an overall coherent project of legal alignment. The goal is for the two domains to productively engage with one another, without either needing to be over-simplified in the process. I would also briefly note here that I don’t think any current definitional boundary between ‘safety’ and ‘alignment’ provides a useful delineation for this discussion, and I will keep using the phrase ‘legal alignment’ to encompass a systems-orientated discussion of the intersection of law and AI, even if that includes consideration of what some would distinguish as ‘AI safety’.276 To start to get to grips with what a systems approach to legal alignment might mean, I think it’s useful to think through existing analogous examples, as opposed to further abstract discussion. To that end, I will briefly consider how a systems approach to AI safety is developing around autonomous driving systems (ADS), albeit this is a brief précis of what could be a much larger comparative study.

Neural networks are critical to ADS, but to narrow questions of safety merely to whether a neural network in an ADS will produce an aligned outcome in a given scenario is to miss the wood and hit the trees. Those who work on the safety of ADS typically treat it as an emergent property of the system rather than of any single component within it.277 When an Uber test vehicle struck and killed Elaine Herzberg in Tempe, Arizona in 2018, the perception network failed to identify her consistently or predict her path, classifying her in turn as a vehicle, an unknown object and a bicycle.278 Nevertheless, the National Transportation Safety Board set that failure in a wider, systemic frame. It traced the crash to Uber’s overall safety culture and to a series of system and organisational choices, including the disabling of the Volvo’s factory automatic emergency braking, a one-second design delay that suppressed the system’s own braking, and an operator-oversight regime too thin to catch that human operators were often distracted by their phones, as was the case in this crash.279 When the National Highway Traffic Safety Administration closed its Tesla Autopilot investigation in 2024, it didn’t focus on the technological capabilities of specific components but on a ‘critical safety gap’ between the system’s capabilities and drivers’ expectations which ‘led to foreseeable misuse and avoidable crashes’.280 When a Cruise vehicle dragged a pedestrian in San Francisco in 2023, the company’s own review highlighted ‘poor leadership’ and ‘a fundamental misapprehension of Cruise’s obligations of accountability and transparency to the government and the public’.281 Partly in anticipation and partly in response, frameworks of authorisation, assurance and accountability that move the focus from the agent to the system and the entity behind it have been building around ADS. In the UK regime (not yet in force) an ‘authorised self-driving entity’ must answer for the vehicle’s behaviour against safety principles published by the state, with a separately licensed operator wherever no one is in charge,282 while the EU type-approves the driving system itself, obliging the manufacturer to run a safety-management system and report its in-service performance.283 Internationally, United Nations Regulation No 157 concerning ‘Automated Lane Keeping Systems’ requires a specific storage system for automated driving and assessment of the manufacturer’s documented safety concept,284 while Germany permits driverless operation in an authority-approved defined operating area with a technical supervisor responsible for specified interventions, but not continuous monitoring.285 The US currently has no comprehensive federal statute, leaning instead on crash reporting and a patchwork of state rules after a proposed national programme was floated in 2025286 and recently withdrawn in 2026.287 The ADS crashes, reports and laws suggest a number of things, however what I note in particular is that contested terms like ‘alignment’ or ‘safety’ are defined through the ongoing dialectic between policy, law and action. This suggests that the goal is not to reach a definition of legal alignment, but to build an evolving system of legal alignment.

I would also suggest that actors attempting to demonstrate legal alignment in a systemic fashion may draw lessons from the attempts to demonstrate ADS safety. A 2025 study by Waymo employees benchmarks the company’s ADS against human drivers and reports fewer injury-causing crashes over 56.7 million driverless miles.288 The human baseline is an estimate built from underreported police-reported data and is difficult to validate.289 Nevertheless, what statistics can show here is bounded by the ‘heavy tail’, in that the rarer the outcome, the more miles a comparison needs.290 The RAND report Driving to Safety found that matching an ADS against the human fatality rate would take ‘hundreds of millions of miles and sometimes hundreds of billions of miles’,291 and that showing one to be a fifth safer than humans would mean driving a hundred vehicles around the clock for roughly five centuries.292 Nevertheless, what gives statistics purchase is that they exist in the context of a bounded operational design domain, so the comparison between human and machine is roughly like-for-like. Where the data are too sparse to be statistically meaningful, particularly in the gravest outcomes, a documented safety case and a minimal-risk fallback are needed.293 I would suggest that law’s vast empire, which we briefly sketched in §1 of this essay, is heavy-tailed in the same way, while currently lacking a settled operational design domain. To benchmark an AI in law as Waymo benchmarks its ADS one would perhaps need a bounded set of questions, an observable outcome with a ground truth, standard metrics like crash-per-mile, a human baseline, and an operating record at some kind of scale. In the ADS analogy LegalBench is perhaps closer to a suite of perception network bench tests than to the crash-per-mile rate of the whole vehicle within its domain. Indeed, LegalBench’s authors caution against treating it as a deployment proxy, warning that performance on it should not be ‘the sole justification for AI deployments’ and is ‘not a substitute for more in-depth and context-specific evaluation efforts’.294 LEXam provides a seemingly more demanding test, evaluating open-ended answers from real law examinations against both reference answers and expected reasoning steps.295 Nevertheless, it is still a test of a model on specific curated questions rather than a record of system performance. Moreover, a recent legal-entailment study suggests that models with the highest benchmark accuracy may also be the least faithful, scoring well by leveraging assumptions formal semantics don’t license.296 Accuracy figures in isolation can select for the failure modes a system-level assurance case exists to rule out. I suggest that the ADS example could illuminate the type of approach to statistical evaluation and benchmarking that could eventually be useful in legal alignment.

Let’s keep the issues, responses and analyses of safety around ADS in mind as we use an imagined future to think about both legal alignment and AGIL. Imagine a British retail bank wires a legal AI system into its databases and applications to make its data protection decisions. Various inputs are wired into the system including access and erasure requests from a web form, breach alerts from the security team and new data uses from the marketing and risk teams. The legal AI assesses this data against the bank’s policies and the UK GDPR and then produces an analysis, perhaps settling a lawful basis for processing, or tracking the 72 hour notification deadline, tens of thousands of times a week.297 The legal AI vendor processes personal data on the bank’s behalf and hosts the system on servers in a Virginia (US) data centre. The bank uses an international data transfer agreement and completes a transfer risk assessment for the transatlantic transfer.298 Suppose the bank also uses customers’ transaction histories to train and run its fraud AI system, a separate system from the legal AI, running on another vendor’s servers and relying on a legitimate interest. Fraud prevention is recognised as a ‘legitimate interest’, separate from the ‘legal obligation’ grounds of the bank’s mandatory anti-money-laundering monitoring.299 The bank treats those histories as ordinary personal data, a classification the legal AI made and recorded in the data protection impact assessment it produced, which was then signed off and filed by compliance staff.300 The systems run untouched for eighteen months, until the UK Supreme Court holds (in this imagined future) that a transaction record can reveal special-category data, such as health or sex life, from a payment to a sexual health clinic, or religious belief from a donation to a place of worship, which a controller cannot lawfully reuse without satisfying a special-category condition. Fraud scores drawn from a history that held such a payment, and the accounts frozen, loans refused and fraud markers filed on the strength of them, may now rest on a reuse of data that breached the UK GDPR. The freezes themselves also fall from the permissive automated-decision regime that governs ordinary personal data into the stricter one reserved for special-category data. A lawful route might have been open, a substantial public interest condition for preventing fraud, had the bank identified it and held the policy document the law requires.301 But the legal AI, having come to a different determination, as a human lawyer may well have also done on the same facts, never raised the question, so none was ever in place. Such payments are common, so the ruling reaches millions of the bank’s customers. For another eight months nobody notices this, and in that window a hardware failure destroys part of the bank’s record of how its decisions were made, logs whose retention had already lapsed, leaving no backup. One of those customers is Jane, whose account the fraud system flagged and froze after reading transactions that included regular payments to a fertility clinic. The frozen account having upended her life, she looked for recourse and so four months ago she complained to the UK ICO but hasn’t heard back, meaning the complaint is now beyond the period of three months within which the regulator is meant to report progress.302 Her record is among those lost in the hardware failure, so the bank can neither properly account for the freeze nor safely undo it, and with no answer from the regulator and nothing to interrogate, she brings proceedings against the bank and both AI vendors.

Taken as a regulative idea for either legal alignment or AGIL, this possible future suggests to me that whether the LLM component of any of the bank’s two AI systems ‘follows’ a given jurisdiction’s law at any one moment (whatever that means) is insufficient to establish whether the deployed systems are accurate, aligned or safe. The data classification and use were defensible when first made, yet the overall system still produced outcomes that became untenable once the law changed, and hard to parse or resolve when viewed only through the lens of whether the LLM component of the respective AI systems knew or followed the law. Set alongside both the AGIL approach I’ve outlined in this essay and the handling of safety around ADS, the scenario suggests that a more systemic approach is possible, one better suited to the many actors, relationships and variables that arise when AI systems are deployed in the real world. While this fact scenario may read as convoluted, it is actually simplified. We didn’t introduce complexities that often arise, such as Jane being a UK citizen and bank customer while resident in Cologne, Germany, potentially bringing the EU GDPR, the BundesdatenschutzgesetzFederal Data Protection Act, and the CJEU’s decision in OT v Vyriausioji tarnybinės etikos komisija into play,303 or consider what it may have meant if one variable the legal AI weighed in classifying the data was non-binding guidance from the UK ICO. There is also the point I alluded to earlier, that legal alignment may be more a question of the intersection of law and ethics within the concept of a legal system. On that view, one might ask whether an AI system aligned in a so-called constitutional fashion took any account of the moral valence of using the fact that Jane had visited a fertility clinic in shaping the bank’s disposition toward her. Perhaps not in the case of everyday banking and fraud, but what if the bank also sold insurance and was running a third AI system assessing Jane’s eligibility for certain insurance products, and the relevant transactions were payments to a drug and alcohol rehabilitation clinic? Whether such a system did weigh that fact, or should have, might be unclear as well, albeit it does get us closer to the question of where the boundary between law and ethics lies in alignment and how one might make it more transparent. Moreover, I would suggest that this brief scenario also speaks to the overarching point being made in §1 of this essay, namely that law’s empire has grown beyond the ability of the legal profession to manage, as many of its leading figures have been saying for some time. I suggest that when you play out such scenarios as regulative ideas, the need for a systems approach to legal AI and legal alignment becomes evident.

6. Conclusion

What is commonly thought of as ‘artificial intelligence’ in our current moment is necessary, but not sufficient, to produce a general legal AI. If one merely looks at the massive volume of complex legal texts, or the fact that lawyers seem to spend much of their time reading, or arguing over words, one might think that an LLM will eventually be able to master legal work, just as Hercules mastered his labours. I’ve argued that this misses something important about what lawyers do on a daily basis, how law works in the world, and what it fundamentally consists of. At a basic level, law is about relations between actors. Legal relations can have stable features, but they can, and frequently do, change significantly, and those changes can recast the past, present and future. Legal relations are both relative to their actors, and the universes in which they’re formed and live. Legal relations can be sufficiently structured such that they can be subject to formal logic, however any such structure must be appropriately aware of its limits and gaps, and what a ‘complete’ legal system looks like. Any ‘general’ legal intelligence worth the appellation must grapple with this reality of the law and how it interconnects with the voluminous palimpsest of overlapping and interconnected norms represented in the legal texts and data that LLMs are particularly adept at navigating and synthesizing. AGIL is constructed on the understanding that, when properly wrought, AI can provide a map to the legal landscape. If law were once a town whose streets an old resident knew like the back of her hand, it is now a vast metropolis, constantly changing and expanding, and interconnected with an even larger world. This is by no means to say that local knowledge has been, or will be, rendered obsolete. There is more to navigating an environment than mapping it and following an algorithmically suggested route in a machine. Rather, the point is that even old locals now regularly use technology to navigate the ever-changing towns and cities their families have lived in for generations. To meet the realities of modernity, we need an artificial general intelligence for law grounded in the law’s standards, history and philosophies.

Metamorphoses of the LawyerFew would describe the experience of being a lawyer in the modern economy as a virtuous pursuit.