Exploring the effectiveness of prompt engineering for legal reasoning tasks

Authors

  • Fangyi Yu Thomson Reuters Labs Author
  • Quarti Li Thomson Reuters Labs Author
  • Frank Schilder Thomson Reuters Labs Author
  • Saida KOHIL Algerian Academy of the Arabic Language Translator

Keywords:

هندسة الأوامر, الاستدلال القانوني, عنقدة البيانات, النماذج اللغوية

Abstract

لقد أفضى استخدام النماذج اللغوية الكبيرة (LLMs) في أساليب الاستدلال الصفري (Zero-shot) أو بتدريب محدود (Few-shot) في ميدان معالجة اللغة الطبيعية، إلى ظهور حقل علمي جديد يعرف بهندسة الأوامر (Prompt Engineering). وقد بينت الدراسات الحديثة أن أوامر سلسلة التفكير (Chain-of-Thought, CoT) تحدث، عند توظيفها، أثرا بينا في مهام عدة كالعمليات الحسابية والاستدلال المبني على الحس المشترك. ويعمد هذا البحث إلى استجلاء أثر تلك الأساليب في الاستدلال القانوني، وذلك عبر تجارب أجريناها على مهمة الاستلزام في مسابقـة COLIEE، وهي مسابقة استخراج المعلومات القانونية والاستلزام القانوني المبنية على اختبار نقابة المحامين اليابانية. وقد قمنا بتقويم أداء النماذج وفق منهج الاستدلال الصفري أو بتدريب محدود، ووفق الضبط الدقيق (Fine-tuning)، مع الشروح أو من دونها، إلى جانب طائفة من طرائق هندسة الأوامر. وتشير نتائجنا إلى أن أوامر سلسلة التفكير، وكذلك الضبط الدقيق المصحوب بعبارات تفسيرية، يحسنان الأداء بدرجات معتبرة. غير أن أفضل النتائج تتحقق عند الاستناد إلى أساليب استدلال قانوني محددة، من بينها منهج IRAC (Issue, Rule, Application, Conclusion)، المبني على عرض المسألة، وبيان القاعدة القانونية، ثم إسقاطها على الواقعة، والانتهاء إلى الحكم. وفضلا على ذلك، اتضح لنا أن التعلم بتدريب محدود، عندما تستمد أمثلته عن طريق عنقدة بيانات التدريب السابقة (Clustering)، يؤدي أداء ثابتا عالي الجودة في مهام الاستلزام ضمن مسابقة COLIEE فخلال سنوات إنجاز البيانات التي خضعت للفحص. وقد أفضت تجاربنا كذلك إلى تحسين أفضل نتائج مسابقة COLIEE لعام 2021 من 0.7037 إلى 0.8025، وتجاوز أحسن الأنظمة المشاركة في عام 2022 بحيث بلغت دقتها 0.789.

 

Downloads

Download data is not yet available.

Author Biographies

  • Fangyi Yu, Thomson Reuters Labs

    Artificial Intelligence Engineer with Machine Learning training from Coursera (June–September 2023), where he developed large language models for behavioral prediction and data analytics. He also serves as an Applied Linguist within the Thomson Reuters Labs team. His expertise includes prompt engineering for large language models (LLMs) and the evaluation of reliability, fairness, and alignment in intelligent AI systems.

    Thomson Reuters Labs
    19 Duncan Street, Toronto, Ontario M5H 3G6, Canada.

           
  • Quarti Li , Thomson Reuters Labs

    Associate Researcher at Thomson Reuters Labs in the field of Artificial Intelligence and its applications in law.
    Thomson Reuters Labs
    3 Times Square, New York, NY 10036, United States.

  • Frank Schilder, Thomson Reuters Labs

    Research Director in the development of intelligent language models and machine learning, specializing in Natural Language Processing (NLP) for intelligent legal solutions and big data analytics. Leads a team of researchers and engineers in developing advanced artificial intelligence technologies for legal language processing and intelligent NLP applications.

    Thomson Reuters Labs
    610 Opperman Drive, Eagan, MN 55123, United States.

           

Downloads

Published

2026-07-19

How to Cite

Exploring the effectiveness of prompt engineering for legal reasoning tasks. (2026). سلسلة أروقة العلوم, 4(6), 09-52. https://corridorsofscience.aala.dz/index.php/corridor/article/view/58