Abstract
Meaning Representation (AMR) emphasizes semantics in a sentence. While training data availability for English is extensive, AMR for non-English languages, such as Indonesian, is limited. Cross-lingual AMR parsing can address such data scarcity, which leverages resources and knowledge from the extensively studied English language. In this research, we developed a cross-lingual AMR parser model for Indonesian sentences using graph pre-training with a multilingual BART model. The vocabulary of the multilingual BART model was trimmed to accelerate the training process without compromising performance. We take advantage of having parallel sentences by concatenating the paired sentences into a single input. We improved the quality and quantity of the available silver dataset for Indonesian AMR parsing by applying two techniques: (1) augmentation using English AMR parsing and machine translation to generate additional silver-quality datasets, and (2) filtration using BLEU metrics and the 1-NN algorithm to filter out low-quality data. Our model achieved a SMATCH score 68.1 on the AMR 2.0 gold dataset. Furthermore, we pushed the performance in the paraphrase detection task using AMR on the WReTE dataset, with an F1 score of 0.758 on the test-set.
| Original language | English |
|---|---|
| Pages (from-to) | 584-603 |
| Number of pages | 20 |
| Journal | International Journal on Electrical Engineering and Informatics |
| Volume | 16 |
| Issue number | 4 |
| DOIs | |
| Publication status | Published - Dec 2024 |
| Externally published | Yes |
Keywords
- Abstract meaning representation
- Cross-lingual parsing
- Dataset augmentation
- Dataset filtration
- Model trimming
Fingerprint
Dive into the research topics of 'Abstract Meaning Representation Parser Development for Cross-lingual Indonesian-English with BART, Input Concatenation, and Dataset Augmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver