Dokument: EXTRACTING MACHINE LEARNING KNOWLEDGE FROM SCHOLARLY PUBLICATIONS : FROM FINE-GRAINED INFORMATION EXTRACTION TO CROSS-SCHEMA GENERALIZATION

Titel:EXTRACTING MACHINE LEARNING KNOWLEDGE FROM SCHOLARLY PUBLICATIONS : FROM FINE-GRAINED INFORMATION EXTRACTION TO CROSS-SCHEMA GENERALIZATION
URL für Lesezeichen:https://docserv.uni-duesseldorf.de/servlets/DocumentServlet?id=74552
URN (NBN):urn:nbn:de:hbz:061-20260928-105631-7
Kollektion:Dissertationen
Sprache:Englisch
Dokumententyp:Wissenschaftliche Abschlussarbeiten » Dissertation
Medientyp:Text
Autor: Otto, Wolfgang [Autor]
Dateien:
[Dateien anzeigen]Adobe PDF
[Details]2,39 MB in einer Datei
[ZIP-Datei erzeugen]
Dateien vom 21.09.2026 / geändert 21.09.2026
Beitragende:Prof. Dr. Dietze, Stefan [Gutachter]
Prof. Dr. Conrad, Stefan [Gutachter]
Stichwörter:Information Extraction, Scholarly Document Processing
Dewey Dezimal-Klassifikation:000 Informatik, Informationswissenschaft, allgemeine Werke » 004 Datenverarbeitung; Informatik
Beschreibungen:The growing number of scholarly publications and the increasing use of machine learning and artificial intelligence across scientific disciplines make it difficult to manually identify, compare, and reuse technical information from the literature. In machine-learning publications, information about models, datasets, tasks, software, baselines, and experimental comparisons is often distributed across the full texts rather than stated in one place.
This thesis is guided by the view that scholarly information extraction (IE) for machine-learning (ML) literature must preserve fine-grained artifact distinctions, links to textual evidence, and awareness of schema differences across annotation and extraction projects. It investigates how such knowledge can be extracted as evidence-linked records of ML artifacts and relations. It further studies how such records can be compared across annotation schemes and extraction settings, and how extraction models can transfer across datasets without collapsing dataset-specific views.
The central goal of the thesis is to develop and evaluate scholarly information extraction methods for machine-learning artifacts and their typed relations in full-text publications. Extraction is treated not only as the detection of text spans, but as the construction of structured records that remain traceable to the sentences from which they were derived. These records distinguish artifact roles that are often merged in broader scientific IE schemas, such as model architectures versus concrete model instances, datasets versus raw-data sources, and named artifacts versus generic artifact mentions. Such distinctions are important because these objects can play different roles in experimental workflows.
The thesis makes five contributions that build on each other, moving from fine-grained extraction resources to cross-schema generalization and interactive use. First, it introduces GSAP-NER, a fine-grained entity schema and manually annotated full-text corpus for ML-related entities in computer science publications, together with baseline models and analyses of annotation and modeling challenges. Second, it studies extraction methods based on large language models by reformulating software-related extraction as single-choice question answering and evaluating retrieval-based prompting in a limited-supervision setting. Third, it extends the entity layer with GSAP-ERE, a full-text entity and relation extraction (ERE) resource that models workflow relations between ML artifacts, including links between models, methods, datasets, tasks, properties, comparisons, and references. Fourth, it investigates cross-schema generalization across scientific ERE benchmarks. Related datasets are first made comparable in a shared evaluation framework that considers both a unified label space and the original dataset-specific label and prediction spaces. Different model and dataset combinations are then trained and evaluated in a schema-preserving multi-dataset setup without giving up the individual schemas of the datasets. Fifth, it presents Methods Miner, a user-facing system that applies the extraction pipeline to paper collections and lets users inspect extracted ML artifacts together with their sentence-level evidence.
The results show that fine-grained extraction of ML artifacts from full-text publications is feasible, but remains limited by ambiguity in conceptual entity types, model references, generic mentions, and relation extraction. They also show that relations are necessary to move from lists of mentioned artifacts to workflow-level representations of experiments. Across scientific ERE benchmarks, the thesis shows that schema alignment enables controlled comparison, but that simple schema unification is not sufficient for robust transfer. A schema-preserving multi-dataset setup is therefore more suitable when datasets share broad scientific targets but differ in annotation density, corpus selection, label granularity, and relation semantics.
Taken together, the results show that artifact granularity, evidence linking, and schema-aware evaluation interact. Fine-grained artifact distinctions improve the usefulness of extracted records only when they remain linked to textual evidence, and schema-aware evaluation is needed to compare such records or transfer extraction models across datasets without removing relevant distinctions. This combination makes extracted records more useful for evaluation, comparison, and interactive exploration, because users can move from aggregated views back to the sentences that support them. At the same time, the thesis identifies important limits of text-only, sentence-level extraction, including cross-sentence relations, document-level coreference, artifact normalization, table and figure evidence, and corpus-wide entity linking. The thesis therefore provides datasets, schemas, models, evaluation setups, and a prototype system as a basis for more reliable extraction and analysis of machine-learning knowledge in scholarly publications.

Die wachsende Zahl wissenschaftlicher Publikationen und der zunehmende Einsatz von maschinellem Lernen (ML) in vielen Fachdisziplinen erschweren es, technische Informationen aus der Literatur manuell zu erfassen, zu vergleichen und wiederzuverwenden. In Publikationen zum maschinellen Lernen sind Informationen zu Modellen, Datensätzen, Tasks, Software, Baselines und experimentellen Vergleichen häufig über den gesamten Volltext verteilt.
Diese Dissertation geht davon aus, dass wissenschaftliche Informationsextraktion für Literatur zum maschinellen Lernen mehr leisten muss, als Textstellen zu markieren. Sie muss erfassen, welche Rolle ML-Artefakte in Experimenten spielen, auf welche Textstellen sich diese Informationen stützen und wie unterschiedliche Annotations- und Extraktionsschemata zueinander stehen. Die Arbeit untersucht daher, wie ML-Artefakte und ihre Beziehungen als strukturierte, beleggestützte Repräsentationen aus Volltextpublikationen extrahiert werden können. Außerdem untersucht sie, wie unterschiedliche Ressourcen und Schemata miteinander vergleichbar gemacht werden können, ohne ihre jeweiligen Perspektiven auf wissenschaftliche Information vollständig zu vereinheitlichen.
Das Ziel der Dissertation ist die Entwicklung und Bewertung von Methoden zur wissenschaftlichen Informationsextraktion (IE) für ML-Artefakte und ihre Beziehungen in Volltextpublikationen. Im Mittelpunkt steht nicht nur, ob ein Artefakt im Text erwähnt wird, sondern auch, welche Art von Artefakt genau gemeint ist und wie es mit anderen Artefakten im Experiment zusammenhängt. Dazu unterscheidet die Arbeit unter anderem zwischen Modellarchitekturen und konkreten Modellinstanzen, zwischen Datensätzen und Rohdatenquellen sowie zwischen benannten Artefakten und generischen Artefakterwähnungen. Diese Unterscheidungen machen sichtbar, welche Funktion ein Artefakt im jeweiligen experimentellen Zusammenhang hat.
Die Dissertation leistet fünf Beiträge, die aufeinander aufbauen und von feingranularen Extraktionsressourcen über schemaübergreifende Generalisierung bis zur interaktiven Nutzung führen. Erstens führt sie GSAP-NER ein, ein feingranulares Entitätsschema und ein manuell annotiertes Volltextkorpus für ML-bezogene Entitäten in Informatikpublikationen, zusammen mit Baseline-Modellen und Analysen zu Annotation und Modellierung. Zweitens untersucht sie die LLM-gestützte Extraktion von Softwareinformationen. Dazu wird der Extraktions-Task als Single-Choice Question Answering reformuliert. Außerdem wird retrieval-basiertes Prompting bei begrenzter Supervision evaluiert. Drittens erweitert sie die Entitätsebene von GSAP-NER zu GSAP-ERE, einer Volltextressource für Entity and Relation Extraction (ERE). Diese Ressource modelliert Workflow-Relationen zwischen ML-Artefakten, darunter Verknüpfungen zwischen Modellen, Methoden, Datensätzen, Tasks, Eigenschaften, Vergleichen und Referenzen. Viertens untersucht sie schemaübergreifende Generalisierung über wissenschaftliche ERE-Benchmarks hinweg. Dazu werden verwandte Datensätze zunächst in einem gemeinsamen Evaluationsrahmen vergleichbar gemacht. Die Evaluation berücksichtigt sowohl einen vereinheitlichten Labelraum als auch die ursprünglichen datensatzspezifischen Label- und Vorhersageräume. Anschließend werden verschiedene Modell- und Datensatzkombinationen in einem schemaerhaltenden Multi-Dataset-Setup trainiert und evaluiert, ohne die individuellen Schemata der Datensätze aufzugeben. Fünftens stellt sie Methods Miner vor, ein nutzerorientiertes System, das die Extraktionspipeline auf Artikelsammlungen anwendet und Nutzern ermöglicht, extrahierte ML-Artefakte zusammen mit ihren Belegen auf Satzebene zu prüfen.
Die Ergebnisse zeigen, dass ML-Artefakte mit einem feingranularen Typsystem aus Volltextpublikationen extrahiert werden können. Diese Extraktion bleibt jedoch durch Mehrdeutigkeit bei konzeptionellen Entitätstypen, Modellverweisen, generischen Erwähnungen und Relationsextraktion begrenzt. Die Ergebnisse zeigen außerdem, wie Relationen ermöglichen, von Listen erwähnter Artefakte zu Darstellungen experimenteller Abläufe zu gelangen. Anhand wissenschaftlicher ERE-Benchmarks zeigt die Dissertation, dass Schema-Abgleich kontrollierte Vergleiche ermöglicht, aber eine einfache Schema-Zusammenführung nicht für robusten Transfer ausreicht. Ein schemaerhaltender Multi-Datensatz-Ansatz ist daher besser geeignet, wenn Datensätze zwar ähnliche wissenschaftliche Ziele verfolgen, sich aber in Annotationsdichte und -stil, Korpusauswahl, Label-Granularität und Relationssemantik unterscheiden. Damit stellt die Dissertation Datensätze, Schemata, Modelle, Evaluationsaufbauten und ein Prototyp-System bereit, die eine nachvollziehbarere Extraktion, Bewertung und Exploration von Wissen zum maschinellen Lernen in wissenschaftlichen Publikationen unterstützen.
Lizenz:Creative Commons Lizenzvertrag
Dieses Werk ist lizenziert unter einer Creative Commons Namensnennung 4.0 International Lizenz
Fachbereich / Einrichtung:Mathematisch- Naturwissenschaftliche Fakultät » WE Informatik
Dokument erstellt am:28.09.2026
Dateien geändert am:28.09.2026
Promotionsantrag am:11.06.2026
Datum der Promotion:02.09.2026
english
Benutzer
Status: Gast
Aktionen