Generative Multilingual Coreference Resolution at CRAC 2026


Jakub Hejman and Ondřej Pražák and Miloslav Konopík
Proceedings of the 2nd Joint Workshop on Computational Approaches to Discourse, Context and Document-Level Inferences and Computational Models of Reference, Anaphora and Coreference (CODI-CRAC 2026) (2026)

PDF

Abstract

Participating again in this year’s edition of the CRAC shared task on coreference resolution, we present our upgraded system with an official uplift of 15.46 percentage points in CoNLL-U score. We incorporated the larger Gemma 3 27B IT model, joint pre-training, headword tagging, more efficient training and inference as well as a sliding window to achieve this result. Our system placed second in the LLM track and third overall with a primary score of 73.83. We reached the highest scores on two datasets. Finally, we compare specialized and general LLM approaches.

Authors

BibTex

@inproceedings{hejman-etal-2026-generative, title = "Generative Multilingual Coreference Resolution at {CRAC} 2026", author = "Hejman, Jakub and Prazak, Ondrej and Konop{\'i}k, Miloslav", editor = "Braud, Chlo{\'e} and Hardmeier, Christian and Ogrodniczuk, Maciej and Loaiciga, Sharid and Zeldes, Amir and Nov{\'a}k, Michal and Li, Chuyuan and Strube, Michael and Li, Junyi Jessy", booktitle = "Proceedings of the 2nd Joint Workshop on Computational Approaches to Discourse, Context and Document-Level Inferences and Computational Models of Reference, Anaphora and Coreference ({CODI}-{CRAC} 2026)", month = jul, year = "2026", address = "San Diego, California, USA", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2026.codi-1.22/", doi = "10.18653/v1/2026.codi-1.22", pages = "162--166", ISBN = "979-8-89176-400-2", abstract = "Participating again in this year{'}s edition of the CRAC shared task on coreference resolution, we present our upgraded system with an official uplift of 15.46 percentage points in CoNLL-U score. We incorporated the larger Gemma 3 27B IT model, joint pre-training, headword tagging, more efficient training and inference as well as a sliding window to achieve this result. Our system placed second in the LLM track and third overall with a primary score of 73.83. We reached the highest scores on two datasets. Finally, we compare specialized and general LLM approaches." }
Back to Top