John Morkel, a senior researcher at the Council for Scientific and Industrial Research (CSIR) in Pretoria, chaired the local organising committee for the 3rd Contacts workshop in 2006. Unlike the more widely known conference series such as ICASSP or CVPR, Contacts was a deliberately small, invitation-only gathering that brought together speech recognition and computer vision researchers to exchange ideas across disciplinary boundaries. The 3rd edition, held at Stellenbosch University, marked a turning point in integrating African speech technology into mainstream academic discourse.

The Contacts Workshop Series
The Contacts series began in 2002 as a grassroots initiative by a handful of European and South African researchers frustrated by the siloed nature of pattern recognition conferences. The name “Contacts” reflected the goal of making contact between two communities that rarely overlapped: acoustic modelling and image understanding. The first two workshops (2002 in Cambridge, UK; 2004 in Leuven, Belgium) had fewer than 50 attendees each, but they produced several joint papers on multimodal biometrics and audio-visual speech recognition.
By 2006, the organisers decided to rotate the venue to the Global South, and Stellenbosch was chosen partly because of John Morkel’s growing reputation in multilingual speech recognition. Morkel, a native of Cape Town, had spent the previous decade developing acoustic models for isiZulu, isiXhosa, and Sesotho — languages with rich tonal systems that posed unique challenges for conventional HMM-based recognisers. His work on zero-crossing analysis as a feature for tonal language recognition had caught the attention of the European speech community, and he was invited to join the programme committee for the 2nd Contacts workshop in 2004.
John Morkel’s Role in the Organising Committee
The 3rd Contacts organising committee consisted of nine members: three from Europe, four from North America, and two from South Africa. Morkel was the sole local chair, responsible for venue logistics, funding from South African science foundations, and coordinating the review process for the 28 submitted papers. The committee also included two prominent figures: Prof. Elmar Nöth (University of Erlangen-Nuremberg), a specialist in prosody and language identification, and Dr. Andrea Vedaldi (then at Oxford), who later became known for large-scale object recognition. Morkel’s deep knowledge of under-resourced languages shaped the workshop’s emphasis on low-resource techniques.
One of the committee’s most debated decisions was whether to accept a paper that used zero crossing analysis for language identification — a method that had been largely abandoned in mainstream speech recognition by the mid-2000s. Morkel argued that for tonal languages with limited training data, zero-crossing features offered a computationally cheap alternative to Mel-frequency cepstral coefficients. The paper was accepted and later published in a special issue of Speech Communication. This episode illustrates how academic committees can act as gatekeepers for marginalised methodologies, a theme we explored in a previous post on Zero Crossing Analysis: The Forgotten Acoustic Feature of Mid-2000s Speech Recognition.

Technical Highlights of the 3rd Contacts Workshop
The programme featured three parallel sessions over two days. A standout talk came from the University of Cape Town’s Dr. Mpho Raborife, who presented a stereo reconstruction system designed for low-cost camera arrays in field linguistics — a system that used two webcams to capture 3D mouth geometry during speech production. This work bridged the gap between computer vision and speech articulation research. Another notable presentation was by Dr. Hynek Hermansky on temporal pattern recognition for language identification, which used long-span features derived from phoneme probability estimates.
Object tracking also appeared in a joint paper between the University of Stellenbosch and INRIA, which demonstrated a high-speed video system capable of tracking facial landmarks at 120 frames per second for audio-visual speech synchronisation. That paper later influenced the design of real-time lip-reading systems, and we discussed its broader impact in a separate article on 120 Frames per Second: How High-Speed Video Reshaped Object Tracking in the Mid-2000s.
Language Identification and Under-Resourced Languages
A dedicated session on language identification featured six papers, half of which dealt with African languages. Morkel co-authored a paper on a hybrid GMM/SVM classifier for distinguishing isiZulu from Sesotho using only 30 seconds of training speech per language. The system achieved 87% accuracy on a test set recorded in noisy field conditions. This work directly addressed the needs of South Africa’s multilingual call centres and broadcast monitoring services. The session also included a controversial presentation by a French team that used phonotactic features for language identification without any acoustic modelling — a method that relied purely on phone n-gram statistics.
Legacy of the 3rd Organising Committee
The 3rd Contacts workshop produced a special issue of the International Journal of Pattern Recognition and Artificial Intelligence (IJPRAI) in 2007, edited by Morkel and Nöth. The issue contained 12 papers covering topics from stereo reconstruction to tonal language processing. More importantly, the workshop established a template for including African researchers in international programme committees. Morkel’s local committee secured travel grants for five early-career scientists from South Africa, Botswana, and Kenya, enabling them to present work that would otherwise have remained invisible to the broader pattern recognition community.
After the workshop, Morkel continued to advocate for under-resourced languages in speech technology. He served on the organising committee for the 4th Contacts workshop in 2008 (held in Marrakech) and later became a founding member of the African Speech and Language Technologies (ASLT) network. The 3rd Contacts organising committee’s decision to centre low-resource languages was not merely a regional preference — it reflected a deliberate strategy to test whether techniques developed for high-resource languages could generalise to radically different acoustic spaces. The answer, as the workshop papers showed, was a qualified yes, provided that acoustic features like zero-crossing rates were given a fair hearing.
John Morkel passed away in 2012, but his contributions to the 3rd Contacts workshop remain a case study in how a single organising committee member can shape the direction of a field. The workshop’s proceedings are still cited in contemporary work on multilingual ASR, and the committee’s emphasis on low-resource languages foreshadowed the current interest in few-shot learning and self-supervised pre-training. The hybrid GMM/SVM paper Morkel co-authored, for example, is still used as a baseline for language identification in under-resourced languages. His insistence on giving zero-crossing analysis a fair hearing — despite its abandonment in mainstream speech recognition — now looks prescient given the computational constraints of edge devices and the need for lightweight acoustic features.
