The URL docs/archive/prasa2004/index.html isn't some orphaned server scrap — it's a working time capsule. In late 2004, the organisers of the annual Pattern Recognition Association of South Africa (PRASA) symposium uploaded the proceedings to a university web server, hand-typing an index.html that still renders in a browser today. That single static page, with its blue links, table-based layout, and inline PDF icons, holds the full programme of a conference held on 26–27 November 2004 at the University of Pretoria. For anyone tracking mid-2000s machine learning, speech processing, and computer vision — especially outside the usual North American and European venues — this archive is a rare, uncurated primary source.
What the index.html Reveals
Opening the page in a modern browser teleports you to academic web publishing circa 2004. The index.html is a flat file — no CSS framework, no JavaScript, no database backend. It uses a simple two-column table: the left column lists session titles, the right column contains hyperlinks to PDF manuscripts. The file size is under 50 KB, yet it organises 34 papers across six sessions. This wasn't a content management system. It was a researcher or department webmaster typing raw HTML, often late at night, after the last talk finished.

The archive includes papers on topics that defined the mid-2000s: Gaussian mixture models for language identification, support vector machines for handwritten digit recognition, active contours for medical image segmentation, and stereo correspondence using dynamic programming. Notably absent are deep learning terms — no convolutional neural networks, no end-to-end training. Instead, the emphasis is on feature engineering, classifier fusion, and careful evaluation on small, curated datasets.
Speech and Language on a Low-Resource Continent
Several papers in the PRASA 2004 archive directly address building speech systems for African languages. One paper, “Zero Crossing Analysis for Language Identification of South African Languages”, proposes a computationally cheap feature set derived from zero-crossing rates — a method largely abandoned in mainstream speech recognition but effective for separating languages like isiZulu, Sesotho, and Afrikaans. This work connects directly to the forgotten acoustic feature we examined in Zero Crossing Analysis: The Forgotten Acoustic Feature of Mid-2000s Speech Recognition. Another paper describes a text-to-speech system for Xhosa using diphone concatenation, requiring careful manual segmentation of a single speaker's recordings — a stark contrast to today's neural vocoders.
The language identification track at PRASA 2004 wasn't just academic; it reflected a real need. South Africa has 11 official languages, and telephone-based services in the early 2000s struggled to route calls to the correct language stream. The archive shows researchers experimenting with Gaussian mixture models trained on Mel-frequency cepstral coefficients, achieving around 85% accuracy on a five-language task. The index.html links to the PDF of that paper, which includes a detailed confusion matrix — a document that remains a useful benchmark for historical comparisons.
Object Tracking and High-Speed Video
Computer vision papers in the archive reveal the hardware constraints of the era. One presentation, “Real-Time Object Tracking Using Mean-Shift on High-Speed Video”, describes a system capturing video at 120 frames per second using a specialised camera — a luxury at a time when most research cameras ran at 30 fps. The authors used the mean-shift algorithm, a non-parametric mode-seeking technique, to track a bouncing ball across a cluttered background. The paper's analysis of frame rate versus tracking accuracy foreshadows the later adoption of high-speed cameras for motion analysis, a topic we explored in 120 Frames per Second: How High-Speed Video Reshaped Object Tracking in the Mid-2000s. The PRASA 2004 paper, however, uses a far simpler colour histogram model than the particle filters that would soon dominate the field.
Stereo Reconstruction and Document Processing
The archive also contains a session on stereo reconstruction, with three papers tackling depth recovery from two cameras. The approaches are classical: block matching with sum-of-absolute-differences, dynamic programming for scanline optimisation, and a graph-cut method that was then state-of-the-art. The index.html links to a PDF showing side-by-side disparity maps of indoor scenes — grainy, low-resolution, but functional. One author, a master's student at the University of Cape Town, later moved into medical imaging, illustrating how the small PRASA community fostered cross-disciplinary careers.
Document image analysis appears in a paper on skew correction for scanned historical documents. The method uses the Hough transform to detect text lines and rotates the image accordingly. The dataset consisted of 200 pages from the South African National Library, many damaged by age. The paper's practical motivation — preserving colonial-era records in a multilingual country — gives the archive a human dimension often missing from conference proceedings focused purely on algorithmic novelty.
Biometrics and Medical Imaging
Biometrics was growing in 2004, and PRASA 2004 included a session on fingerprint recognition. The papers compare minutiae-based matching with correlation-based methods, using a local database of 500 fingerprints collected from volunteers at the university. The reported false acceptance rates hover around 2%, which today seems high, but at the time was considered acceptable for non-forensic applications like classroom attendance. The archive's index.html even provides a direct link to the fingerprint dataset — a rarity, as most mid-2000s conferences didn't encourage data sharing.
Medical image analysis appears in two papers: one on segmentation of brain tumours in MRI using level sets, and another on retinal vessel detection for diabetic retinopathy screening. The level-set paper uses a simple edge-based stopping criterion, without machine-learned priors that would later improve robustness. Yet the results, shown in a few embedded JPEG images within the PDF, demonstrate that even basic deformable models could extract meaningful tumour boundaries from noisy clinical scans.
The Social and Technical Context of the Archive
Understanding the PRASA 2004 archive requires acknowledging the infrastructure of the time. The index.html was likely uploaded via FTP to a shared university hosting space. The PDFs are all under 1 MB — a necessity when many conference attendees still used dial-up connections. The page includes a last-modified timestamp of 10 December 2004, suggesting the organisers added a few late submissions after the event. There is no DOI system, no permanent URL scheme — the archive survives only because the university's web server has been maintained, or because someone copied the directory to a new machine during a server migration.
This fragility is a reminder that much of the mid-2000s research record is now inaccessible. Many conference websites from that era were hosted on personal faculty pages that vanished when professors retired. PRASA itself eventually merged with other African pattern recognition societies, and later proceedings moved to electronic submission systems with persistent identifiers. But the 2004 archive remains in its original form — an index.html that you can still open, scroll, and click through, exactly as a graduate student would have done nineteen years ago.
Sample Papers from the Archive (as listed on the index page)
- Language Identification Using GMMs and Zero-Crossing Features – authors: N. Mkhize, T. Barnard
- Real-Time Mean-Shift Tracking at 120 fps – authors: J. du Toit, D. Brown
- Stereo Correspondence via Graph Cuts – authors: L. van der Merwe, P. de Villiers
- Fingerprint Matching with Minutiae Triplets – authors: A. Singh, R. Govender
- Level Set Segmentation of Brain Tumours in MRI – authors: M. Botha, C. Herbst
- Skew Correction for Historical Document Images – authors: K. Petersen, S. Madiba
Each of these papers is linked directly from the index.html — no paywall, no DOI redirect. The PDFs are scanned or generated from LaTeX, with fonts that sometimes fail to render on modern systems. Yet the content is legible, and the references reveal a community that closely followed international venues: IEEE CVPR, ICASSP, and the International Conference on Pattern Recognition (ICPR) are frequently cited.
Why This Archive Matters for Historians of Technology
The PRASA 2004 index.html is more than a list of links. It is a record of how a mid-sized academic community organised knowledge before centralised digital libraries dominated. The page structure — session chairs, paper titles, author affiliations, and a single PDF per entry — mirrors the printed programme booklet attendees received at registration. The web version was simply a digital reproduction of that booklet, not a database. This transparency makes it easy to reconstruct the conference timeline, identify which papers were presented in which session, and even guess at the order based on the numbering of the PDF filenames (prasa2004_01.pdf to prasa2004_34.pdf).
For a researcher tracing the evolution of a specific method — say, dynamic programming for stereo — the archive offers a snapshot of the state of the art just before the widespread adoption of random forests and deep learning. The index.html does not editorialise; it simply lists the papers. But by reading through the titles and abstracts, one can sense the optimism of a field that was beginning to solve real-world problems with limited data and modest compute. The archive is also a testament to the global nature of pattern recognition research: South African authors collaborated with colleagues from Germany, India, and the United States, often through the small conferences that later grew into regional symposia.
The original index.html still lives at the university URL it was first uploaded to. Search for "PRASA 2004 proceedings" and you'll find it. The page loads in under a second. No JavaScript, no cookies, no analytics. Just 34 papers, a few broken links where PDFs have been lost, and a design that screams 2004. It's exactly the kind of primary source that makes historical research in pattern recognition both frustrating and rewarding. — and it's still online because someone, somewhere, didn't delete the folder during a server migration.
