Generative AI has become remarkably good at answering questions. ChatGPT and Claude can explain complex topics, summarize documents and produce convincing legal-sounding text within seconds.
But there is an important distinction:
Sounding like a lawyer is not the same as reliably answering a legal question.
This becomes particularly important when AI is used professionally in areas such as employment law, contracts, corporate law or taxation.
The fundamental problem: hallucinations
Large language models generate answers based on statistical patterns. They do not inherently determine whether a legal conclusion is correct.
This creates a well-known problem: hallucinations.
An answer can sound perfectly plausible while containing an incorrect interpretation, an outdated rule or even a source that does not support the conclusion.
For casual questions, this may be acceptable.
For professional legal decisions, it is not.
Adding RAG does not automatically solve this
A common approach is to connect ChatGPT, Claude or another general-purpose model to a database containing laws, court decisions or other legal documents.
This technique is usually called Retrieval-Augmented Generation (RAG).
RAG is useful. It can provide the model with relevant documents and substantially improve the quality of its answers.
But there is an important misconception:
RAG alone does not turn a general-purpose LLM into a reliable legal reasoning system.
Retrieving a relevant document is only the beginning.
The system still has to determine which provisions actually apply, understand their relationship, distinguish statutory law from case law and other sources, deal with exceptions and conflicting information, and verify that the final conclusion is supported by the retrieved material.
A chatbot interface + an LLM + a vector database may therefore look like a legal AI system without having solved the difficult part of legal AI.
Retrieval is not reasoning
Consider a relatively simple employment-law question.
An employee becomes ill during a notice period. What happens to the termination date?
Finding the relevant Swiss legal provision is not particularly difficult.
But answering the actual question may require determining the applicable blocking period, when it starts, whether and for how long the notice period is interrupted, how the remaining notice period is calculated and whether the resulting termination date needs to be moved to the end of a month.
The challenge is therefore not simply:
“Can the AI find the law?”
It is:
“Can the system reliably reason from the relevant law to the correct answer?”
That distinction becomes increasingly important as questions become more complex.
This is what Jurilo was built for
Jurilo was not created by simply placing a chatbot interface around a general-purpose LLM.
It was developed specifically for Swiss legal questions.
The system combines curated Swiss legal sources with Jurilo's Legal Graph™ and multi-step reasoning architecture. A question can trigger more than 25 agentic operations before the final answer is produced.
The objective is not merely to generate an answer.
It is to identify the relevant legal framework, retrieve authoritative sources, reason across them and provide a conclusion whose derivation can be understood and checked.
Jurilo currently focuses on Swiss law and uses authoritative sources including Swiss legislation and Federal Supreme Court decisions.
The benchmark should be correctness, not eloquence
This is perhaps the most important point when evaluating legal AI.
Almost every modern LLM can produce an impressive-looking legal answer.
That is no longer particularly difficult.
The meaningful questions are:
Is the applicable law correctly identified?
Are the sources authoritative and relevant?
Does the cited source actually support the statement?
Are exceptions and dependencies considered?
Is the reasoning internally consistent?
Can the user understand how the conclusion was reached?
Does the system know when the available evidence is insufficient?
This is where legal AI systems should be compared.
Not by which chatbot produces the nicest prose.
Even the general-purpose AI providers recognize the risk
The providers of general-purpose AI systems themselves place limitations and warnings around high-impact professional uses.
OpenAI's policies restrict certain high-impact decisions in legal and other sensitive domains when they are made without appropriate human involvement. Anthropic likewise provides usage restrictions and guidance around high-impact applications.
That does not mean ChatGPT or Claude are poor products. Quite the opposite: they are extraordinarily capable general-purpose AI systems.
But general-purpose capability and specialized legal reliability are two different engineering objectives.
Building another interface around those models does not remove that distinction.
Legal AI needs to be engineered for trust
The next generation of professional AI applications will not be defined simply by which foundation model they use.
The foundation models will increasingly become infrastructure.
The differentiation will be in everything built around them: domain architecture, authoritative data, retrieval, reasoning, verification, traceability and continuous testing.
For legal AI, this is particularly important.
The question is no longer whether AI can answer a legal question.
It clearly can.
The question is whether you can trust the answer when the decision actually matters.
That is the problem Jurilo was built to solve.

