Data cleaning and spam filtering methods in the TAWOS database

Szabó, Márk, Kovács, Ádám, Kusper, Gábor (2026) Data cleaning and spam filtering methods in the TAWOS database Annales Mathematicae et Informaticae. 63. pp. 132-148.

[thumbnail of 132_148.pdf] pdf
132_148.pdf

Download (700kB) [error in script]
Hivatalos webcím (URL): https://doi.org/10.33039/ami.2026.06.004

Absztrakt (kivonat)

The TAWOS dataset is a large relational database of issue-tracking data mined from open-source projects. While it is valuable for empirical software engineering, the issue texts also contain spam, placeholders, and off-topic noise that can distort downstream analytics and fine-tuning tasks. We present a four-stage filtering pipeline for cleaning TAWOS and focus on a validity scoring step that assigns each issue a 0–100 score from its title and description. We compare a deterministic rule-based classifier, OwnMetrics, against four local small language models (Llama-3.1-8B, Mistral-7B, Phi-3.5-mini, and Gemma-3-4B) prompted for JSON scores. Evaluation first uses a near-balanced labeled benchmark of 947 GitHub issues (481 spam/noise and 466 legitimate issues), collected from moderator-locked spam issues and legitimate issues from popular repositories. On this GitHub-based threshold-selection benchmark, OwnMetrics obtains the highest accuracy, 91.0%, at threshold 75, while the best LLM configurations reach 89.2% (Gemma-3-4B) and 89.1% (Mistral-7B). A separate manually labeled in-domain validation on 300 Jira/TAWOS issues (39 spam/noise and 261 non-spam) provides an in-domain sanity check consistent with the benchmark results, with OwnMetrics giving the strongest spam-class F1 among the selected configurations. On a 10,000-item sample, OwnMetrics yields zero parsing failures and an average execution time of 2.28 ms per item, whereas the LLMs require seconds per item and Meta-Llama-3.1-8B exhibits failure rates above 54%. Applying the full filtering pipeline to the processed TAWOS issue corpus flags 16.2% of records for filtering. The results show that a lightweight domain-tailored scorer can be both accurate and robust for large-scale issue cleaning.

Mű típusa: Folyóiratcikk - Journal article
Szerző:
Szerző neve
Email
MTMT azonosító
ORCID azonosító
Közreműködés
Szabó, Márk
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Kovács, Ádám
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Kusper, Gábor
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
NEM RÉSZLETEZETT
Szerző
Kapcsolódó URL-ek:
Kulcsszavak: software repositories, issue tracking, data cleaning, spam filtering, large language models
Nyelv: angol
Kötetszám: 63.
DOI azonosító: 10.33039/ami.2026.06.004
Felhasználó: Tibor Gál
Dátum: 20 Júl 2026 07:31
Utolsó módosítás: 20 Júl 2026 07:31
URI: http://publikacio.uni-eszterhazy.hu/id/eprint/9300
Műveletek (bejelentkezés szükséges)
Tétel nézet Tétel nézet