Publications
Sexism has pervasive negative effects on both individuals and society. This paper presents a generalizable BERT-based approach to identifying and classifying the source intent of sexism across different social network channels. This approach focuses on individual models trained on the text of tweets and then applied to both Meme (image) and Video data using OCR and annotations respectively. The identification model performed well across all channels and the classification model performed well on both Tweets and Memes. This research suggests that a single model, fine-tuned on one media type can be effectively applied to multiple media types with minimal data preprocessing required.
Accurate Species Distribution Modelling (SDM) is essential for biodiversity conservation, however the limited and spatially biased nature of Presence-Absence (PA) data poses a challenge. In contrast, Presence-Only (PO) datasets are abundant but lack explicit absence records. This paper examines a two step deep learning approach to combining both PO and PA data to generate an SDM. In the first step, the model was trained on a larger PO dataset, and in the second the model was then tuned on a smaller PA dataset. Results indicate that pre-training with PO data improved the performance by 7% when subsequently fine-tuned with PA data, as measured by the samples-averaged F1-score. This approach demonstrates the potential of combining diverse data types to create more reliable species distribution models for plant biodiversity conservation.
- Rawlings, D., & Chopard, T. (2024). Exploring biodiversity: A multi-model approach to multi-label plant species prediction. In 25th Working Notes of the Conference and Labs of the Evaluation Forum, CLEF 2024 (pp. 2188-2200). CEUR Workshop Proceedings.
The research aims to develop a multi-label classification method for predicting plant species based on environmental factors such as satellite images, climate data, and other environmental variables. Utilizing a dataset of plant surveys in Europe, the study addresses challenges such as collating image, tabular and time series data from a variety of sources, multi-modal learning, with variable label counts. The approach employs dimensionality reduction via PCA, and an ensemble of machine learning models used to predict both species pseudo-probabilities and also the counts of species present in a given area. The potential contributions of this research include advancing our understanding of ecological systems, informing conservation efforts, and promoting the preservation of plant diversity - all essential components of a comprehensive approach to safeguarding the natural world.