Szabó, Márk, Kovács, Ádám, Kusper, Gábor (2026) Data cleaning and spam filtering methods in the TAWOS database Annales Mathematicae et Informaticae. 63. pp. 132-148.
|
pdf
132_148.pdf Download (700kB) [error in script] |
Absztrakt (kivonat)
The TAWOS dataset is a large relational database of issue-tracking data mined from open-source projects. While it is valuable for empirical software engineering, the issue texts also contain spam, placeholders, and off-topic noise that can distort downstream analytics and fine-tuning tasks. We present a four-stage filtering pipeline for cleaning TAWOS and focus on a validity scoring step that assigns each issue a 0–100 score from its title and description. We compare a deterministic rule-based classifier, OwnMetrics, against four local small language models (Llama-3.1-8B, Mistral-7B, Phi-3.5-mini, and Gemma-3-4B) prompted for JSON scores. Evaluation first uses a near-balanced labeled benchmark of 947 GitHub issues (481 spam/noise and 466 legitimate issues), collected from moderator-locked spam issues and legitimate issues from popular repositories. On this GitHub-based threshold-selection benchmark, OwnMetrics obtains the highest accuracy, 91.0%, at threshold 75, while the best LLM configurations reach 89.2% (Gemma-3-4B) and 89.1% (Mistral-7B). A separate manually labeled in-domain validation on 300 Jira/TAWOS issues (39 spam/noise and 261 non-spam) provides an in-domain sanity check consistent with the benchmark results, with OwnMetrics giving the strongest spam-class F1 among the selected configurations. On a 10,000-item sample, OwnMetrics yields zero parsing failures and an average execution time of 2.28 ms per item, whereas the LLMs require seconds per item and Meta-Llama-3.1-8B exhibits failure rates above 54%. Applying the full filtering pipeline to the processed TAWOS issue corpus flags 16.2% of records for filtering. The results show that a lightweight domain-tailored scorer can be both accurate and robust for large-scale issue cleaning.
| Mű típusa: | Folyóiratcikk - Journal article |
|---|---|
| Szerző: | Szerző neve Email MTMT azonosító ORCID azonosító Közreműködés Szabó, Márk NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző Kovács, Ádám NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző Kusper, Gábor NEM RÉSZLETEZETT NEM RÉSZLETEZETT NEM RÉSZLETEZETT Szerző |
| Kapcsolódó URL-ek: | |
| Kulcsszavak: | software repositories, issue tracking, data cleaning, spam filtering, large language models |
| Nyelv: | angol |
| Kötetszám: | 63. |
| DOI azonosító: | 10.33039/ami.2026.06.004 |
| Felhasználó: | Tibor Gál |
| Dátum: | 20 Júl 2026 07:31 |
| Utolsó módosítás: | 20 Júl 2026 07:31 |
| URI: | http://publikacio.uni-eszterhazy.hu/id/eprint/9300 |
![]() |
Tétel nézet |
