Skip to main navigation Skip to search Skip to main content

Knowledge and error profile
: a context-based meta-analysis framework for synergistic error sampling in active learning

  • Qurat Ul Ain Shaheen

Student thesis: Doctoral Thesis (PhD)

Abstract

Active learning is a popular machine learning paradigm that allows a learning algorithm to select the most informative instances for training. The goal is to reduce data labelling cost and, consequently data consumption by eliminating redundant data instances. Classification error serves as the primary signal for identifying informative instances in most active learning strategies. In the active learning literature, however, it is typically treated as an isolated, unidimensional quantity, disconnected from cognitive theories of human error.

This thesis focuses on pool-based batch-mode active learning for binary classification with structured data. We propose that treating classification error as a semantically rich concept, grounded in Smithson's (1989) taxonomy of ignorance from cognitive and social science, offers two key benefits for active learning. The first is that such an approach can explicitly leverage synergistic error interactions for more informative query selection than conventional methods such as query by bagging. The second is that it enables a more nuanced understanding of error dynamics in a dataset.

To test this hypothesis, we propose a novel context-based meta-analysis framework, called Knowledge and Error Profile (KEP), that tracks three types of errors in a learning algorithm's knowledge base: uncertainty (incomplete information), distortion (misrepresentation of information) and absence (lack of information).

We empirically evaluate the framework on three publicly available real world datasets through three analyses. First, a performance comparison shows that KEP outperforms query by bagging and random sampling. Second, an analysis of three error-type synergies in KEP reveals that there is no universal error synergy across datasets, rather the optimal error synergy depends on the dataset. Third, a sensitivity analysis of our framework shows how to tune the hyperparameters for best results. Based on these empirical analyses, we identified best practices for applying KEP in practice. Overall, this work promotes a new perspective on classification error and opens new research directions for query strategies based on error semantics.
Date of Award1 Dec 2026
Original languageEnglish
Awarding Institution
  • University of St Andrews
SupervisorJuliana Bowles (Supervisor)

Keywords

  • Active learning
  • Knowledge representation
  • Error analysis
  • Qualitative analysis
  • Error aggregation

Access Status

  • Full text embargoed until
  • 15 Aug 2028

Cite this

'