Skip to main navigation Skip to search Skip to main content

Data Segmentation and High Dimensional Time Series Analysis

  • Dom M Owens

Student thesis: Doctoral ThesisDoctor of Philosophy (PhD)

Abstract

Time series analysis, the study of time-ordered data, is an historic and accomplished subfield
of statistics. A great number of methods have been proposed for the analysis of time series
data, and these have found use in many application areas. Often, however, these methods
correspond to models that do not account for two properties commonly found in the data. These
are non-stationarity, wherein the joint distribution of the underlying process is not constant, and
high dimensionality, where the number of concurrent series is possibly larger than the sample
size. This thesis proposes new methods for the statistical analysis of data with either or both of
these properties. We preface the thesis with a review of relevant literature.
First, in Chapter 3, we describe a data segmentation procedure for high dimensional data
which follow a regression model, where the parameters are piecewise-constant with respect to
time. In two steps, the method first compares parameter estimates over a moving window to
detect changes, then uses local refinements minimising a comparative loss for the final location
estimates. We prove that the method consistently detects and locates all changes at the minimax-
optimal rate (up to logarithmic factors) under Gaussianity, and is consistent under dependence
and heavy tails, and when changes are multiscale. The computational cost is small relative to
competing algorithms.
In Chapter 4, we describe a segmentation procedure for multiple time series which follow a
piecewise-stationary vector autoregression (VAR) model. The method uses a moving sum detector
to look for changes in the expectation of estimating functions. We prove this is consistent for
the number and locations of changes when the dimension is fixed. A series of methodological
extensions are proposed so that the algorithm may be used in less idealised settings, for example
with changes which are multiscale or only detectable with a local inspection parameter.
Chapter 5 describes an extension of the model and method from Chapter 4, where we treat
high-dimensional time series as observations from a dynamic factor model, allowing the latent
factors to follow a piecewise-stationary VAR. By applying the proposed method to estimated
principal components, we show that we consistently segment the data. We give a similar method
for segmenting a piecewise stationary factor-augmented regression, and combine this with
methods for forecasting under structural breaks, giving a comprehensive method for diffusion
index forecasting for non-stationary data.
Finally in Chapter 6 we discuss the factor-adjusted VAR model, designed for high dimen-
sional time series which exhibit strong serial and cross-sectional correlations, as well as sparse
idiosyncratic structure. We give an overview of the model, particularly estimation procedures
for the common and idiosyncratic components and of the implicit network structures, as well
as forecasting methods. A software package is described, with visualisation methods and tools
for the data-driven selection of tuning parameters, and we extensively study the computational
properties of the method.
We end the thesis with a discussion and directions for future work.
Date of Award23 Jan 2024
Original languageEnglish
Awarding Institution
  • University of Bristol
SupervisorHaeran Cho (Supervisor)

Keywords

  • Statistics

Cite this

'