Skip to main content
GalsenAI

Research and data

GalsenAILab

What the community builds, and what comes out of it. Projects on one side, the corpora they produce on the other.

Projects

What the community builds

Work supported by GalsenAI and collaborative initiatives it participates in.

  • AfriQA

    Collaborative

    Creation of an open-access evaluation dataset for answering questions in African languages.

    • Question-réponse
    • Langues africaines
    • Corpus
    Active
  • Masakhane

    Collaborative

    Research community working on NLP for African languages.

    • NLP
    • Traduction automatique
    • Recherche ouverte
    Active
  • Waxal

    GalsenAI

    A voice recognition project for Wolof.

    • Reconnaissance vocale
    • Langues africaines
    • Corpus
    Active
  • Adama

    GalsenAI

    A Wolof text-to-speech model.

    • Synthèse vocale
    • Langues africaines
    Active
  • Aya

    Collaborative

    An open-science initiative by Cohere For AI to build a state-of-the-art multilingual language model.

    • Modèles de langue
    • Multilingue
    • Science ouverte
    Active
  • A machine translation project between French and Wolof.

    • Traduction automatique
    • Langues africaines
    Active

Datasets

What comes out of it

Senegalese languages remain poorly endowed with exploitable data. This section gathers corpora produced by the community and those it recommends, detailing their domain, languages, volume, and license.

13 datasets listed

  • Synthèse vocale

    Cleaned and denoised version of the female voice from the Wolof TTS corpus. Annotations were revised to remove special characters, emojis and foreign characters, and clips judged to be of insufficient quality were discarded. Represents 18 h 41 min of usable speech.

    • Wolof
    Volume
    19,947 audio clips · 18 h 41 min
    License
    CC BY 4.0

    Open on Hugging Face

  • Reconnaissance vocale

    Automatic speech recognition dataset for Wolof, built by bringing together four existing corpora and their transcriptions. Designed for training ASR models, it is the largest transcribed audio collection in Wolof published by the community.

    • Wolof
    Volume
    35,075 audio clips · 6.5 GB
    License
    Apache 2.0

    Open on Hugging Face

  • Synthèse vocale

    Speech synthesis corpus recorded by two native Wolof speakers, one male and one female voice, each having read more than 20,000 sentences. Collected by Baamtu Datamation as part of the AI4D African Language Program, it is the reference basis for speech synthesis in Wolof.

    • Wolof
    Volume
    40,010 audio clips · 4.6 GB
    License
    CC BY 4.0

    Open on Hugging Face

  • Traduction automatique

    Aggregation of the various Wolof-French translation sources available within the community, brought together in a single format with traceability of each pair's origin. This is the largest Wolof-French parallel corpus published by GalsenAI.

    • Wolof
    • Français
    Volume
    98,345 sentence pairs
    License
    License not specified by the producer , check the terms of reuse before use

    Open on Hugging Face

  • Traduction automatique

    Small English-Wolof parallel corpus, designed for rapid evaluation and for bootstrapping translation models into English.

    • Wolof
    • Anglais
    Volume
    7,405 sentence pairs
    License
    License not specified by the producer , check the terms of reuse before use

    Open on Hugging Face

  • Traduction automatique

    Reference evaluation set for machine translation, covering 200 languages including Wolof. Because the same sentences are translated into every language, it makes it possible to compare a Wolof model against results obtained on better-resourced languages.

    • Wolof
    • Français
    • Anglais
    • Multilingue
    Volume
    200 languages
    License
    CC BY-SA 4.0

    Open on Hugging Face

  • Traduction automatique

    French-Wolof parallel corpus aligned sentence by sentence, each pair retaining a reference to its original source. Used to train the community translation models.

    • Wolof
    • Français
    Volume
    17,777 sentence pairs
    License
    License not specified by the producer , check the terms of reuse before use

    Open on Hugging Face

  • Reconnaissance vocale

    Transcribed speech corpus devoted to agriculture in the three most widely spoken languages of Senegal. The recordings bring together farmers, agricultural advisors and agrifood business managers, through interactive radio programmes, focus groups, voice messages and interviews. Project funded by the Lacuna Fund.

    • Wolof
    • Pulaar
    • Sérère
    Volume
    125 h of transcribed speech, 35 h of them verified
    License
    CC BY 4.0

    Open on Zenodo

  • Traduction automatique

    Parallel translation corpus in the news domain, covering 21 African languages including Wolof, with French and English as pivot languages. Also produced by the Masakhane collective.

    • Wolof
    • Français
    • Anglais
    • Multilingue
    Volume
    21 African languages
    License
    CC BY-NC 4.0

    Open on Hugging Face

  • Reconnaissance d'entités

    Named entity recognition corpus covering 20 African languages, including Wolof. Produced by the Masakhane collective, to which GalsenAI members contribute. Annotations cover people, places, organisations and dates.

    • Wolof
    • Multilingue
    Volume
    20 African languages
    License
    AFL 3.0

    Open on Hugging Face

  • Classification audio

    Voice command dataset covering Wolof, Pulaar and Serer, annotated by keyword. Each audio clip is paired with one of the words in the target vocabulary, making it a resource suited to keyword spotting and to voice interfaces in national languages.

    • Wolof
    • Pulaar
    • Sérère
    Volume
    Between 10,000 and 100,000 clips
    License
    CreativeML OpenRAIL-M

    Open on Hugging Face

  • Reconnaissance vocale

    Correction of the Wolof portion of the WAXAL corpus published by Google, whose audio files and transcriptions suffered from a misalignment that made the resource unusable. The community regenerated the transcriptions and realigned the whole set, making the corpus usable again.

    • Wolof
    Volume
    1,042 audio clips · 6.1 GB
    License
    License not specified by the producer , check the terms of reuse before use

    Open on Hugging Face

  • Corpus textuel

    Monolingual text corpus in Wolof, intended for the pre-training and adaptation of language models. Serves as the basis for the Wolof language modelling work carried out within the community.

    • Wolof
    Volume
    52,706 sentences
    License
    License not specified by the producer , check the terms of reuse before use

    Open on Hugging Face

A project to share?

If you are working on Senegalese languages, on an open corpus, or on a model that the community could reuse, write to us.