Karin van Wyk and the Contacts 3 Organising Committee: A Mid-2000s Milestone

When the third Contacts workshop convened in Stellenbosch in November 2006, the organising committee faced a challenge that would define the event: how to bring together researchers working on speech and language technology for languages with virtually no digital resources. Karin van Wyk, then a senior lecturer at North-West University’s School of Computer Science and Information Systems, was one of the key figures on that committee. Her expertise in acoustic phonetics and her hands-on work with Setswana speech corpora made her an indispensable bridge between the signal-processing community and the field linguists who had begun to document under-resourced African languages.

Three researchers reviewing a workshop schedule at a whiteboard in a university conference room

The Contacts Workshop Series

The Contacts series—short for “Computational Techniques for African Speech and Language”—emerged in the early 2000s as a response to the growing awareness that mainstream speech recognition and synthesis systems ignored the vast majority of the world’s languages. The first two workshops (2003 in Pretoria, 2004 in Nairobi) had established a small but dedicated community. By the time Contacts 3 was being planned, the field had matured enough to demand a more structured organising committee, one that could balance academic rigour with practical outreach to language activists and government agencies.

The committee comprised seven members, each bringing a distinct skill set. Van Wyk represented the speech-processing wing, alongside colleagues from the University of Cape Town and the CSIR. Other members covered computational linguistics, fieldwork methodology, and evaluation metrics. The chair was Prof. Marelie Davel, a pioneer in Afrikaans speech recognition, but van Wyk’s role was particularly crucial because she had recently completed the first phonetically balanced speech corpus for Setswana—a language spoken by over five million people in South Africa and Botswana.

Karin van Wyk: A Profile

Van Wyk’s academic journey began at the University of Pretoria, where she earned her master’s degree in electronic engineering with a thesis on formant tracking in tonal languages. In 2004, she moved to North-West University and began collaborating with the Human Language Technology Research Group. Her work on the Setswana corpus involved recording 200 speakers across three dialect regions, manually annotating phonetic boundaries, and designing a small-vocabulary recogniser that could handle the language’s complex noun-class system.

What set van Wyk apart from many of her contemporaries was her insistence on grounding every algorithm in real linguistic data. She often argued that acoustic models trained on English or Mandarin would fail for languages like Setswana because of differences in vowel space, tone, and co-articulation patterns. This perspective, though now widely accepted, was still controversial in the mid-2000s, when many researchers believed that universal feature sets could be adapted with minimal tuning. As discussed in our earlier post on Zero Crossing Analysis: The Forgotten Acoustic Feature of Mid-2000s Speech Recognition, the reliance on simple acoustic features was a practical necessity when working with limited training data—a principle van Wyk applied directly in her Setswana recogniser.

The Organising Committee for Contacts 3

The committee’s work began in early 2005. Van Wyk was responsible for the technical programme, which meant she had to solicit papers, assign reviewers, and design the oral and poster sessions. She also coordinated a special track on “Data Collection and Annotation for Under-Resourced Languages”—a topic she knew intimately. The committee met three times in person: once in Johannesburg, once in Stellenbosch, and once via a rudimentary video-conferencing system that often dropped audio. Van Wyk later recalled in an interview that the biggest difficulty was not technical but social: convincing African language speakers that their speech data would not be exploited or misused.

The final programme for Contacts 3 included 18 full papers, 5 system demonstrations, and 2 invited talks. One of the highlights was a live demonstration of a Setswana text-to-speech system built by van Wyk’s student, which used diphone concatenation—a technique that was already fading in mainstream TTS but remained the most reliable approach for languages with sparse data. The workshop also featured a roundtable on ethical data sharing, which led to the creation of the African Speech Data Commons, a repository that still exists today.

A computer screen showing waveform and spectrogram with labelled phonetic segments

Challenges of Under-Resourced Language Technology in the Mid-2000s

To appreciate the significance of van Wyk’s role on the organising committee, one must understand the technological landscape of 2005–2006. Open-source speech toolkits like HTK and Sphinx existed but required significant expertise to adapt. The standard approach—training hidden Markov models on hundreds of hours of transcribed speech—was simply impossible for languages with fewer than 10 hours of annotated data. Van Wyk and her committee colleagues promoted alternative methods:

  • Cross-lingual bootstrapping: using acoustic models from a related language (e.g., Zulu to bootstrap for Xhosa) and then adapting them with small amounts of target-language data.
  • Active learning: selecting the most informative utterances for human transcription, reducing annotation effort by up to 40%.
  • Articulatory feature extraction: using knowledge of tongue and lip positions to generate acoustic features that generalise better across languages.

These techniques were presented in several papers at Contacts 3, and van Wyk’s own contribution—a paper on “Acoustic-Phonetic Similarity Metrics for Setswana Vowels”—became one of the most cited works from the workshop. She also organised a tutorial on using the Praat software for phonetic annotation, which attracted over 30 participants, many of whom were linguists with no programming background.

Legacy of Contacts 3

The workshop’s impact extended well beyond the two days of presentations. The proceedings, published as a special issue of the South African Computer Journal, served as a de facto handbook for anyone starting work on an under-resourced language. Van Wyk’s organising committee experience also shaped her later career: she went on to co-chair the 2008 Workshop on African Language Technology (AfLaT) and served on the programme committee for Interspeech’s special session on language diversity.

Perhaps the most enduring outcome was the informal network that Contacts 3 solidified. Van Wyk maintained contact with several participants, and together they launched a collaborative project to build a multilingual speech corpus for ten South African languages. That corpus, completed in 2009, remains a primary resource for researchers today. Van Wyk’s own work on the Setswana speech corpus, which she first presented at Contacts 3, later became the foundation for the first commercial text-to-speech system for the language, deployed in a public information kiosk at the Potchefstroom campus in 2010.