De grootste kennisbank van het HBO

Inspiratie op jouw vakgebied

Vrij toegankelijk

Deel deze publicatie

Identifying SME needs in the Zeeland region using unsupervised text analysis

Open access

Rechten:

Identifying SME needs in the Zeeland region using unsupervised text analysis

Open access

Rechten:

Samenvatting

This research originated from the need to transform unstructured internship assignment descriptions from the OnStage dataset into structured insights that support a better understanding of SME challenges in the Zeeland region. The objective was to contribute to the desired future situation in which the KCOI can use these insights to improve its understanding of regional SME needs. The main research question that guided this project was:

“How can internship assignment descriptions from the OnStage database be analysed using a generic and reusable data science approach to identify and group challenges faced by SMEs in the Zeeland region?”

To address this, an unsupervised NLP pipeline was developed by combining Sentence Transformer embeddings, UMAP dimensionality reduction, K-Means clustering and BERTopic topic representation using c-TF-IDF.

The sub-questions supported the realisation of this objective by guiding the development of the full pipeline, from understanding stakeholder expectations to data characteristics to preprocessing decisions, modelling choices, evaluation and interpretation. Together, they prove that internship descriptions can be transformed into structured clusters. However, these clusters should be treated as exploratory indicators, rather than fixed categories of SME needs.

The dataset initially consisted of 28 columns, with only the internship description one being relevant for analysis. Almost half of the internship descriptions were missing. The remaining ones vary in size from a single word to more than 3000 words. After preprocessing, the final dataset contained 7570 descriptions.

The clustering approach produced more than five distinct topics, thus addressing the first BOSC, which required at least five recurring categories. However, these clusters do not directly represent SME challenges, but they function as indicators towards underlying organization needs. To satisfy the second BOSC, three visualizations that show both the separation and interaction between clusters have been reviewed by the stakeholder, as presented in Qualitative Data Inspection.

Regarding the third BOSC, the methodology and pipeline were reviewed with Mischa Beckers and refined accordingly, resulting in a structured workflow. While this improves transparency and supports potential reuse, the approach remains at least partly tailored to the OnStage dataset.

Model evaluation against the DMSC confirmed that all required thresholds are met, with the final model exceeding expectations on Silhouette Score. Refinements to the pipeline were done iteratively, with the trade-off between technicality and interpretability in mind. Validation with Ageeth van Maldegem confirmed that the clusters were meaningful for exploratory purposes.

In conclusion, this study demonstrated that extracting SME challenges can be partly achieved through an embedding-based unsupervised NLP pipeline that transforms unstructured internship descriptions into structured clusters. The approach successfully supports the KCOI in gaining exploratory understanding of SME-related challenges in Zeeland, therefore contributing to their plan of helping regional organizations. However, while the BOSC are largely met in terms of number of clusters, visualisations and methodology transparency, the resulted outputs should not be interpreted as definitive SME challenges.

Toon meer
Organisatie
Opleiding
Afdeling
PartnerHZ University of Applied Sciences, Middelburg
Datum2026-06-29
Type
TaalEngels

Op de HBO Kennisbank vind je publicaties van 26 hogescholen

De grootste kennisbank van het HBO

Inspiratie op jouw vakgebied

Vrij toegankelijk