A light method for data generation: a combination of Markov Chains and Word Embeddings.
dc.contributor.author | Martínez García, Eva | |
dc.contributor.author | Nogales Moyano, Alberto | |
dc.contributor.author | Morales Escudero, Javier | |
dc.contributor.author | García Tejedor, Álvaro José | |
dc.date.accessioned | 2021-06-16T08:19:29Z | |
dc.date.available | 2021-06-16T08:19:29Z | |
dc.date.issued | 2020 | |
dc.description.abstract | Most of the current state-of-the-art Natural Language Processing (NLP) techniques are highly data-dependent. A significant amount of data is required for their training, and in some scenarios data is scarce. We present a hybrid method to generate new sentences for augmenting the training data. Our approach takes advantage of the combination of Markov Chains and word embeddings to produce high-quality data similar to an initial dataset. In contrast to other neural-based generative methods, it does not need a high amount of training data. Results show how our approach can generate useful data for NLP tools. In particular, we validate our approach by building Transformer-based Language Models using data from three different domains in the context of enriching general purpose chatbots. | spa |
dc.description.extent | 1,74 MB | spa |
dc.identifier.doi | 10.26342/2020-64-10 | spa |
dc.identifier.issn | 1135-5948 | spa |
dc.identifier.uri | http://hdl.handle.net/10641/2327 | |
dc.language.iso | eng | spa |
dc.publisher | Procesamiento del Lenguaje Natural | spa |
dc.relation.publisherversion | http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6199 | spa |
dc.rights | Atribución-NoComercial-SinDerivadas 3.0 España | * |
dc.rights.accessRights | open access | spa |
dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/3.0/es/ | * |
dc.subject | Generation | spa |
dc.subject | Hybrid | spa |
dc.subject | Markov Chains | spa |
dc.subject | Embeddings | spa |
dc.subject | Similarity | spa |
dc.title | A light method for data generation: a combination of Markov Chains and Word Embeddings. | spa |
dc.title.alternative | Un método ligero de generación de datos: combinación entre Cadenas de Markov y Word Embeddings. | spa |
dc.type | journal article | spa |
dc.type.hasVersion | AM | spa |
dspace.entity.type | Publication | |
relation.isAuthorOfPublication | 26d8c0ac-90ac-4130-925c-7b36638291c7 | |
relation.isAuthorOfPublication | b8353a15-5990-41e6-88db-6f9489ed2635 | |
relation.isAuthorOfPublication | f3703882-8d88-448b-871c-b450bcd59001 | |
relation.isAuthorOfPublication.latestForDiscovery | 26d8c0ac-90ac-4130-925c-7b36638291c7 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- 6199-5608-1-PB.pdf
- Size:
- 1.75 MB
- Format:
- Adobe Portable Document Format
- Description: