Species sampling processes have long provided the reference framework for modelling random discrete distributions and exchangeable sequences. Data arising from distinct but related sources, however, require a broader notion of probabilistic invariance, making partial exchangeability a natural choice. Numerous models for partially exchangeable data, known as dependent nonparametric priors, have been proposed, including hierarchical, nested and additive processes used in statistics and machine learning. Still, a unifying theory is lacking and key questions about their underlying learning mechanisms remain unanswered. We fill this gap by introducing multivariate species sampling models, a new class encompassing most existing finite- and infinite-dimensional dependent processes. They are characterised by the partially exchangeable partition probability function encoding their multigroup clustering structure. We establish their core distributional properties and analyse their dependence structure, showing that borrowing of information across groups is governed by the probability of ties across groups. This clarifies their learning mechanisms and gives a principled rationale for the previously unexplained correlation structure observed in existing models. Beyond a cohesive theoretical foundation, the approach is a constructive tool for building new models and opens directions for richer forms of dependence.
Main reference: B. Franzolini, A. Lijoi, I. Prünster and G. Rebaudo (2026+). Multivariate Species Sampling Processes. arXiv:2503.24004.