Abstract
Amachine learning task usually starts with the collection and annotation of data or with aprovided curated dataset that needs to be understood. Before any learning takes place, the first
concern is data quality as this affects any subsequent tasks and inferences [67]. This has led to
the development of tools for addressing transparency and accountability of data [136? ]. An equally
important, and often overlooked concern in supervised learning is the quality of labels. For example,
labels are expected to not be ideal in situations where the data is harvested directly from the web
[36, 127]. In general, this is a result of annotations not being carried out by domain experts.
In the first part of the thesis, in Chapter 2, we introduce a framework for inspecting and categorising
weak supervision settings. It can be used for understanding the flexibility and implications of
deciding on an annotation process, for describing a dataset, or for deciding on suitable machine learning
algorithms. In Chapters 3 and 4 we explore in more detail one of the settings: label proportions
and extensions of it.
One of the properties analysed in the weak supervision framework is concerned with the symmetry
of the label noise. If the noise satisfies this property then -under mild conditions- it does not bias
the learning process. In the second part of the thesis we first study the implications symmetry has
on learning in more depth, in Chapter 5. In Chapter 6 we build on this to create tools for inspecting
the quality of labels in a supervised dataset. These take the form of hypothesis tests for assessing
whether the dataset has been subject to asymmetric label noise or not.We present hypothesis tests to
check whether a given dataset of instance-label pairs has been corrupted with asymmetric label noise,
as opposed to symmetric label noise. The outcome of these tests can then be used in conjunction
with other information to assess further steps. What we present is designed to be performed after
data collection and annotation to offer a quality measure with respect to label noise. If the quality is
deemed poor then the practitioner could resort to: (1) a modified data labelling procedure (e.g., active
learning in the presence of noise), (2) seek methods to make the training robust (e.g., algorithms for
learning from noisy labels), or (3) drop the dataset altogether.
The third part of the thesis looks into machine learning interpretability as a tool for inspecting
models. The ability to extract complementary explanations offers the practitioner a new medium
of assessing model performance beyond accuracy on a left-out dataset. This information would
be beneficial in determining the limitations of the model and subsequently its applicability. In
Chapter 7 we look at local surrogates and in Chapter 8 we look at counterfactuals. Local surrogates
and counterfactuals are tools that can be used to inspect a model, assess what the model has learned,
or if the model is compliant with restrictions such as being fair across sub-populations. In Chapter 7
we provide an overview of existing research in the field and show that approaches under the same
category of local-surrogate explainers have different objectives and capture different information
from the black-box. We then discuss a different view of surrogates that begins from surrogates at
the global level being reduced to surrogates at the local level. In Chapter 8 we study counterfactual
explanations, we discuss a new direction of research and propose an algorithm for this.
| Date of Award | 27 Sept 2022 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Niall Twomey (Supervisor) & Raul Santos-Rodriguez (Supervisor) |
Cite this
- Standard