AfriQA
CollaborativeCreation of an open-access evaluation dataset for answering questions in African languages.
- Question-réponse
- Langues africaines
- Corpus
Research and data
What the community builds, and what comes out of it. Projects on one side, the corpora they produce on the other.
Projects
Work supported by GalsenAI and collaborative initiatives it participates in.
Creation of an open-access evaluation dataset for answering questions in African languages.
Research community working on NLP for African languages.
A voice recognition project for Wolof.
A Wolof text-to-speech model.
An open-science initiative by Cohere For AI to build a state-of-the-art multilingual language model.
A machine translation project between French and Wolof.
Datasets
Senegalese languages remain poorly endowed with exploitable data. This section gathers corpora produced by the community and those it recommends, detailing their domain, languages, volume, and license.
13 datasets listed
Synthèse vocale
Cleaned and denoised version of the female voice from the Wolof TTS corpus. Annotations were revised to remove special characters, emojis and foreign characters, and clips judged to be of insufficient quality were discarded. Represents 18 h 41 min of usable speech.
Open on Hugging Face
Reconnaissance vocale
Automatic speech recognition dataset for Wolof, built by bringing together four existing corpora and their transcriptions. Designed for training ASR models, it is the largest transcribed audio collection in Wolof published by the community.
Open on Hugging Face
Synthèse vocale
Speech synthesis corpus recorded by two native Wolof speakers, one male and one female voice, each having read more than 20,000 sentences. Collected by Baamtu Datamation as part of the AI4D African Language Program, it is the reference basis for speech synthesis in Wolof.
Open on Hugging Face
Traduction automatique
Aggregation of the various Wolof-French translation sources available within the community, brought together in a single format with traceability of each pair's origin. This is the largest Wolof-French parallel corpus published by GalsenAI.
Open on Hugging Face
Traduction automatique
Small English-Wolof parallel corpus, designed for rapid evaluation and for bootstrapping translation models into English.
Open on Hugging Face
Traduction automatique
Reference evaluation set for machine translation, covering 200 languages including Wolof. Because the same sentences are translated into every language, it makes it possible to compare a Wolof model against results obtained on better-resourced languages.
Open on Hugging Face
Traduction automatique
French-Wolof parallel corpus aligned sentence by sentence, each pair retaining a reference to its original source. Used to train the community translation models.
Open on Hugging Face
Reconnaissance vocale
Transcribed speech corpus devoted to agriculture in the three most widely spoken languages of Senegal. The recordings bring together farmers, agricultural advisors and agrifood business managers, through interactive radio programmes, focus groups, voice messages and interviews. Project funded by the Lacuna Fund.
Open on Zenodo
Traduction automatique
Parallel translation corpus in the news domain, covering 21 African languages including Wolof, with French and English as pivot languages. Also produced by the Masakhane collective.
Open on Hugging Face
Reconnaissance d'entités
Named entity recognition corpus covering 20 African languages, including Wolof. Produced by the Masakhane collective, to which GalsenAI members contribute. Annotations cover people, places, organisations and dates.
Open on Hugging Face
Classification audio
Voice command dataset covering Wolof, Pulaar and Serer, annotated by keyword. Each audio clip is paired with one of the words in the target vocabulary, making it a resource suited to keyword spotting and to voice interfaces in national languages.
Open on Hugging Face
Reconnaissance vocale
Correction of the Wolof portion of the WAXAL corpus published by Google, whose audio files and transcriptions suffered from a misalignment that made the resource unusable. The community regenerated the transcriptions and realigned the whole set, making the corpus usable again.
Open on Hugging Face
Corpus textuel
Monolingual text corpus in Wolof, intended for the pre-training and adaptation of language models. Serves as the basis for the Wolof language modelling work carried out within the community.
Open on Hugging Face
If you are working on Senegalese languages, on an open corpus, or on a model that the community could reuse, write to us.