AI legal research should be evaluated against the law, languages and tasks a firm actually handles. A fluent answer, a Swiss address or a list of citations does not establish that the research is complete, current or suitable for a client matter.
For Swiss practices, start with a source-coverage assessment and a controlled comparison with the existing research process. Product descriptions are leads for that assessment, not independent evidence of accuracy.
What the research establishes
Magesh and colleagues' preregistered evaluation, published in the Journal of Empirical Legal Studies in 2025, tested commercial legal research systems in 2024. Retrieval-augmented generation, which supplies documents to a model, reduced hallucinations relative to the general-purpose comparator but did not eliminate them. The study also found differences in responsiveness and accuracy between systems.
This is evidence about the tested versions and US legal questions. It does not rank current Swiss products, establish their error rates, or prove that Swiss citations are more vulnerable than US citations. Use it to challenge a claim of error-free research and to design local tests.
A complementary 2026 randomised study by Schwarcz and colleagues found quality and productivity gains when upper-level law students used specified research or reasoning tools for bounded legal tasks. It did not test an entire Swiss client engagement. Together, the studies support testing both possible benefits and remaining errors.
Check coverage before comparing answers
Ask suppliers for the actual source inventory, update process, subscription restrictions and exclusions. Check these statements against sample searches. A reference to “Swiss law” is too broad to establish coverage of a particular canton's decisions, a commentary series or historical statutory versions.
Jurisdiction and authority
Evidence to request: Included legislation, courts and commentary
Local check: Retrieve known primary and secondary sources
Language
Evidence to request: Supported search and source languages
Local check: Test German, French and Italian where relevant
Currency
Evidence to request: Update dates and version handling
Local check: Compare amended law and later decisions
Answer support
Evidence to request: Document and passage references
Local check: Check whether each passage supports the proposition
Data handling
Evidence to request: Contract, subprocessors, access and retention
Local check: Map the actual configured service and data flows
Cross-language retrieval needs its own tests. Ask questions in one language whose relevant authority is in another. Inspect the original source and distinguish a translation from authoritative wording. Do not infer this capability from the languages available in the interface.
Verify propositions, not just citation existence
A real judgment can be irrelevant, misquoted, superseded or used outside its procedural context. Review the proposition, supporting passage, date, jurisdiction and subsequent treatment. Compare important omissions against an independently prepared research set.
A model's self-check can help generate questions for review. It is not independent verification. Where the answer turns on competing interpretations, assess the reasoning and counterarguments as well as the cited authorities.
Assess secrecy and data protection separately
The Swiss Bar Association's February 2025 AI guidance, published in German, French and Italian only, describes professional considerations and possible arrangements for AI use. It does not endorse a vendor or make Swiss hosting a general safe harbour. Examine supplier access, storage, subprocessors, contracts and the particular matter against professional secrecy and applicable data protection requirements.
The FDPIC confirms that the current FADP applies directly to AI-supported data processing. Under Article 5 letter c FADP, sensitive personal data are data on religious, ideological, political or trade union views or activities; health, the intimate sphere or racial or ethnic origin; genetic data; biometric data that uniquely identify a person; administrative and criminal proceedings or sanctions; and social assistance measures. Legal files often contain such data, for example in a criminal matter or a personal injury claim involving health data, but a matter is not sensitive merely because it is legal work or a civil dispute: its contents determine the classification. A data protection impact assessment is required where processing may lead to a high risk, and Article 22 FADP names extensive processing of sensitive personal data as such a case. A firm that will use a research tool to process large volumes of such files therefore generally needs to carry out that assessment, as controller, before the processing starts; the presence of AI alone does not decide the question.
Design a pilot that can reveal failure
Ada Studio proposes the following evaluation approach. It is a planning aid, not a validated minimum sample or a guarantee of legal compliance.
Research-tool evaluation
0/8Select the sample size and trial duration according to the variety and consequences of the work. A few successful demonstrations cannot establish performance across a practice. Record failures and the work needed to correct them before deciding whether to expand use.
Key Takeaway
Get in touch to discuss an evaluation for your practice areas and languages.