https://ujsds.ppj.unp.ac.id/index.php/ujsds/issue/feed UNP Journal of Statistics and Data Science 2026-05-31T12:23:01+00:00 Open Journal Systems UNP Journal of Statistics and Data Science https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/474 Random Forest Algorithm Implementation for Air Quality Classification in DKI Jakarta Based on ISPU 2026-04-27T01:42:41+00:00 Khairanisa Salsabila khairanisa2004@gmail.com Tessy Octavia Mukhti tessyoctaviam@fmipa.unp.ac.id <p><em>Air quality is an essential factor that has a direct impact on human health. High concentrations of air pollutants have the potential to cause various health impacts, across short-term and long-term horizons. This study aims to classify air quality in DKI Jakarta using the Air Pollution Standard Index (ISPU) data via the random forest algorithm. The dataset covers a timeframe from 2021 to 2025 and includes air pollutant parameters, namely PM10 and PM2.5 particulate matter, carbon monoxide (CO), nitrogen dioxide (NO2), sulfur dioxide (SO2), dan ozone (O3). The research method employs a supervised learning approach, in which the data are stratified and evakuated through the implementation of K-Fold Cross Validation (k = 10) to ensure objective and stable model performance. Model performance was measured using Accuracy, Precision, Recall, and F1-Score metrics, along with Confusion Matrix and Feature Importance analyses. It can be seen from the results that the Random Forest model can classify air quality categories with excellent performance, reaching 100% Accuracy on training data and 98.44% on testing data. The Confusion Matrix analysis indicates that most data in each air quality are correctly classified. Furthermore, the Feature Importance analysis reveals PM2.5 that is most influential parameter in determining air quality categories. Therefore, this study indicates that the Random Forest algorithm proves effective for air quality classificati and can function as a decision-support tool for air pollution control and management in DKI Jakarta.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Khairanisa Salsabila, Tessy Octavia Mukhti https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/473 A Predicting the Future: A Forecast of Bukittinggi's Original Local Revenue from 1996 to 2024 2026-05-25T03:40:26+00:00 Fedisha Elfiri Fedisha fedisha01@gmail.com Fadhilah Fitri fadhilahfitri@fmipa.unp.ac.id Zilrahmi zilrahmi@fmipa.unp.ac.id <p><em>This study aims to accurately forecast Bukittinggi City’s Original Local Revenue (PAD) for the period from 2025 to 2029 by applying the ARIMA (Autoregressive Integrated Moving Average) modeling approach. Secondary data were obtained from the Bukittinggi City BPS website, covering annual revenue figures from 1996 to 2024. The analysis begins with a thorough exploratory examination to identify trends and key fluctuations in PAD, including the marked decline during the COVID-19 pandemic in 2020, which significantly impacted revenue patterns. To ensure reliable forecasts, the time series data undergo Box-Cox transformations and differencing to achieve stationarity required for ARIMA modeling. Various diagnostic tests, such as the Ljung-Box test for residual autocorrelation and the Kolmogorov-Smirnov test for normality of residuals, guide the appropriate selection of the most suitable model. The ARIMA (1,1,0) model emerges as the best fit, yielding robust and reliable forecasting results. The projections estimate an average annual PAD value of approximately 138.167billion Rupiah throughout the 2025 to 2029 timeframe. However, forecast uncertainty increases over time, indicated by wider confidence intervals in the later years. These important findings highlight the critical importance of thorough fiscal planning and diversification of revenue sources for Bukittinggi City to withstand potential economic shocks and uncertainties caused by both internal and external factors. The research strongly underscores the necessity for adaptive and resilient financial strategies to sustain steady local revenue growth amid varying and challenging economic conditions. This comprehensive study provides valuable insights to support informed policy-making and effective revenue management by local government authorities.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Fedisha Elfiri Fedisha, Fadhilah Fitri, Zilrahmi https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/475 Factors Affecting Turnover Intention: A Survival Analysis Approach with the Stratified Cox Model 2026-05-05T09:27:05+00:00 Reihan Dani Eka Saputra rehantanjung199@gmail.com Tessy Octavia Mukhti tessyoctaviam@fmipa.unp.ac.id <p><em>Employee turnover remains a critical management challenge in Indonesia. This research investigates the predictors of turnover intention—age, gender, profession, and transportation mode—using a survival analysis approach. Utilizing a real-world dataset of 1,129 observations, initial diagnostics indicated that the 'profession' variable violated the proportional hazard (PH) assumption (p &lt; 0.05), necessitating a Stratified Cox Proportional Hazard model. We developed and compared two model configurations—with and without interaction terms—using the Akaike Information Criterion (AIC). While the non-interaction model proved most optimal for overall prediction (AIC: 5124.104), the interaction model revealed nuanced dynamics across professional strata. Key findings indicate that age generally increases turnover risk by 6.3% per year (HR: 1.063), and walking to work provides a protective effect, reducing risk by 13.6% (HR: 0.864) compared to bus usage. However, professional context significantly modulates these effects: in the 'Manage' stratum, age serves as a stabilizer (HR: 0.822), whereas male teachers face a risk 200.8% higher than their female counterparts (HR: 3.008). Furthermore, car usage in the 'Consult' stratum leads to a dramatic 423.5% increase in turnover risk (HR: 5.235). These results underscore the necessity of strata-specific retention strategies that prioritize workplace accessibility and demographic inclusivity. This study provides a robust data-driven framework for organizations to maintain workforce stability amidst the evolving labor landscape in Indonesia.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Reihan Dani Eka Saputra, Tessy Octavia Mukhti https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/478 Hybrid LBFA-Based Feature Selection for Improving Machine Learning Classification Performance in Heart Disease Prediction 2026-05-12T09:00:59+00:00 Hana Azizah azizah.hanasa@gmail.com Eni Sumarminingsih eni_stat@ub.ac.id Adji Achmad Rinaldo Fernandes fernandes@ub.ac.id <p>Feature selection and feature engineering play an important role in improving the predictive performance of machine learning models, particularly when dealing with real-world datasets that often contain redundant variables and imbalanced class distributions. However, many feature augmentation approaches are implemented without a consistent preprocessing structure, which may lead to information leakage or suboptimal feature representations. To address this issue, this study proposes a hybrid classification framework based on Logit-Based Feature Augmentation (LBFA) integrated with a structured preprocessing pipeline. In the proposed framework, the original predictors are first filtered using CatBoost-based feature selection to identify the ten most relevant variables. Feature augmentation is then performed through two complementary transformations, namely the LOGIT transformation and the Log Density Ratio (LDR). Each transformation is constructed through a dedicated preprocessing branch to ensure methodological consistency. One-hot encoding is applied specifically for the construction of LOGIT features, while numerical standardization is used only for estimating LDR features. The augmented features are subsequently combined with the selected original predictors to form several input configurations used to train gradient boosting classifiers, including XGBoost, LightGBM, and CatBoost. Experiments are conducted using the Heart Disease dataset, and model performance is evaluated using accuracy, precision, recall, specificity, and F1-score. The best performance is achieved by the LightGBM + LOGIT + LDR model with an accuracy of 0.8493, precision of 0.2531, recall of 0.8189, specificity of 0.8512, and F1-score of 0.9874, demonstrating that integrating feature selection with consistent feature augmentation can improve predictive performance in imbalanced classification problems.</p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Hana Azizah, Eni Sumarminingsih, Adji Achmad Rinaldo Fernandes https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/476 Stock Price Forecasting of PT Bank Rakyat Indonesia (Persero) Tbk Using the Support Vector Regression Method 2026-05-06T04:46:29+00:00 Widya Febriani wfebriani978@gmail.com Dony Permana donypermana@fmipa.unp.ac.id <p>Peramalan harga saham merupakan aktivitas penting di pasar modal karena pergerakan harga saham cenderung nonlinier dan volatil dari waktu ke waktu. PT Bank Rakyat Indonesia (Persero) Tbk (BBRI) adalah saham unggulan dengan likuiditas tinggi dan fundamental yang kuat, menjadikannya subjek yang tepat untuk penelitian peramalan. Studi ini bertujuan untuk memprediksi harga saham BBRI menggunakan metode Support Vector Regression (SVR), yang dikenal karena kemampuannya untuk memodelkan hubungan nonlinier dan meminimalkan overfitting. Data yang digunakan terdiri dari harga penutupan harian BBRI dari Januari 2020 hingga Desember 2024. Sebelum pemodelan, data dinormalisasi menggunakan metode Min-Max dan dibagi menjadi set pelatihan dan pengujian dengan rasio 80:20. Model dasar awal menggunakan SVR dengan kernel linier. Model kemudian dioptimalkan menggunakan kernel Radial Basis Function (RBF) melalui Grid Search Optimization yang dikombinasikan dengan validasi silang deret waktu untuk menentukan kombinasi parameter terbaik. Parameter optimal dipilih berdasarkan Root Mean Square Error (RMSE) terendah. Hasil penelitian menunjukkan bahwa model SVR RBF mengungguli model linier dalam menangkap pola nonlinier harga saham BBRI. Selama pengujian, model yang dioptimalkan mencapai RMSE sebesar 0,022054, yang menunjukkan akurasi prediksi yang tinggi. Model SVR yang dioptimalkan kemudian digunakan untuk memprediksi harga saham untuk periode berikutnya dan menunjukkan pergerakan harga yang relatif stabil namun dinamis. Secara keseluruhan, temuan ini menegaskan bahwa metode SVR efektif dan dapat diandalkan untuk prediksi harga saham dan dapat menjadi referensi berharga bagi investor dan penelitian keuangan di masa mendatang.</p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Widya Febriani Widya, Dony Permana https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/479 Regresi Data Panel dengan Kesalahan Standar Driscoll-Kraay: Analisis Kejahatan dan Indikator Sosial Ekonomi di Sumatera Barat (2017–2024) 2026-05-07T03:40:26+00:00 Andini Diva Luthfiyah andinidiva2004@gmail.com Dhio Ervandi dhioervandi@gmail.com Tessy Octavia Mukhti tessyoctaviam@fmipa.unp.ac.id <p>Perilaku kriminal merupakan masalah sosial yang kompleks yang mengancam ketertiban umum dan menghambat pembangunan daerah. Di Indonesia, tingkat kejahatan bervariasi antar provinsi dan dipengaruhi oleh berbagai faktor sosioekonomi dan struktural. Di Provinsi Sumatera Barat, fluktuasi risiko kejahatan dari waktu ke waktu menyoroti perlunya analisis yang lebih mendalam terhadap faktor-faktor penentunya. Memahami faktor-faktor ini sangat penting bagi pemerintah untuk merumuskan kebijakan pencegahan kejahatan yang efektif dan tepat sasaran. Penelitian ini bertujuan untuk menganalisis faktor-faktor penentu risiko kejahatan di Provinsi Sumatera Barat menggunakan data panel dari tahun 2017 hingga 2024, yang mencakup 19 kabupaten dan kota, sehingga memungkinkan evaluasi yang lebih kuat dan komprehensif terhadap variasi temporal dan cross-sectional. Variabel yang dianalisis meliputi tingkat pengangguran terbuka, tingkat kemiskinan, persentase pemuda yang tidak bekerja, bersekolah, atau mengikuti pelatihan (NEET), serta pandemi COVID-19 sebagai variabel dummy. Analisis regresi data panel digunakan, dan hasilnya menunjukkan bahwa model yang paling tepat adalah Model Efek Acak (REM). Temuan menunjukkan bahwa tingkat pengangguran terbuka dan variabel pandemi memiliki pengaruh yang signifikan terhadap risiko kejahatan pada tingkat signifikansi 5%, sedangkan tingkat kemiskinan signifikan pada tingkat 10%. Hasil ini memberikan wawasan berharga bagi pembuat kebijakan dalam menangani akar penyebab kejahatan di Sumatera Barat melalui penciptaan lapangan kerja, pengentasan kemiskinan, dan kesiapan menghadapi situasi krisis.</p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Andini Diva Luthfiyah, Dhio Ervandi, Tessy Octavia Mukhti https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/482 Spatial Analysis of Open Unemployment Rate in West Java Province Using the Spatial Autoregressive Model 2026-04-27T02:28:55+00:00 Zulfadly Harman Harahap zulfadlyhrphrp@gmail.com Tessy Octavia Mukhti tessyoctaviam@fmipa.unp.ac.id <p><em><span dir="auto" style="vertical-align: inherit;"><span dir="auto" style="vertical-align: inherit;">Pengangguran tetap menjadi isu sosial-ekonomi utama di Provinsi Jawa Barat. Tingkat Pengangguran Terbuka (KPR) tidak hanya dipengaruhi oleh faktor regional internal tetapi juga oleh kondisi daerah tetangga, yang menunjukkan adanya saling ketergantungan spasial. Studi ini bertujuan untuk menganalisis pola spasial KPR di Jawa Barat dan mengidentifikasi faktor-faktor yang mempengaruhinya menggunakan pendekatan Autoregresif Spasial (SAR). Studi ini menggunakan data sekunder lintas sektoral dari seluruh kabupaten dan kota di Jawa Barat untuk tahun 2023. Hasil uji Moran's I menunjukkan autokorelasi spasial positif, yang menunjukkan bahwa daerah dengan KPR tinggi biasanya dikelilingi oleh daerah dengan tingkat pengangguran yang sama tingginya. Berdasarkan uji Pengali Lagrange, model SAR dipilih. Hasil estimasi menunjukkan bahwa tingkat pertumbuhan penduduk dan pengeluaran pemerintah secara signifikan mempengaruhi KPR. Selain itu, koefisien lag spasial positif dan signifikan, yang menegaskan adanya efek spillover spasial. Temuan ini menyoroti pentingnya memasukkan perspektif spasial dalam merumuskan kebijakan ketenagakerjaan daerah.</span></span></em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Zulfadly Harman Harahap, Tessy Octavia Mukhti https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/484 Sentiment Analysis of Public Opinion on Rupiah Redenomination on Twitter Using Naive Bayes Classification 2026-05-24T09:10:12+00:00 FIGO RAHMATULLAH figorahmatullah868@gmail.com Dila Sari figorahmatullah868@gmail.com Rahmat Kurniawan raehmatkurniawan@gmail.com Fadhilah Fitri fadhilahfitri@fmipa.unp.ac.id <p>Studi ini meneliti opini publik tentang kebijakan redenominasi Rupiah melalui analisis sentimen data Twitter. Redenominasi mengacu pada penyederhanaan denominasi mata uang tanpa mengubah nilai riilnya, sebuah kebijakan yang sering memicu beragam respons publik karena kekhawatiran seperti persepsi inflasi dan ilusi uang. Di era digital, Twitter (saat ini X) berfungsi sebagai platform utama untuk ekspresi publik secara real-time, menghasilkan sejumlah besar data tekstual tidak terstruktur yang cocok untuk analisis. Tujuan penelitian ini adalah untuk mengklasifikasikan sentimen publik terhadap kebijakan redenominasi Rupiah ke dalam kategori positif, negatif, dan netral menggunakan Naive Bayes Classifier, serta untuk mengevaluasi kinerja model. Dataset terdiri dari tweet berbahasa Indonesia yang dikumpulkan melalui API Twitter menggunakan kata kunci yang terkait dengan redenominasi. Pemrosesan data melibatkan beberapa tahapan, termasuk pembersihan data, pelabelan manual, pra-pemrosesan teks (case folding, tokenisasi, penghapusan stopword, dan stemming), dan ekstraksi fitur menggunakan Term Frequency–Inverse Document Frequency (TF–IDF). Hasil klasifikasi dievaluasi menggunakan matriks kebingungan. Klasifikasi Naive Bayes mencapai akurasi sekitar 74,84% dan presisi 80%, menunjukkan bahwa model tersebut berkinerja memadai dalam mengidentifikasi pola sentimen. Temuan menunjukkan bahwa sentimen netral mendominasi diskusi, menunjukkan bahwa sebagian besar pengguna cenderung memberikan opini informatif atau observasional daripada dukungan atau penentangan yang kuat. Hasil ini diharapkan dapat memberikan wawasan bagi para pembuat kebijakan, khususnya Bank Indonesia dan pemerintah, mengenai penerimaan publik terhadap kebijakan redenominasi, sekaligus berkontribusi pada pengembangan penelitian analisis sentimen pada data media sosial Indonesia.</p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 FIGO RAHMATULLAH, Dila Sari, Rahmat Kurniawan, Fadhilah Fitri https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/485 Application of the Cox Proportional Hazards Model to Analyze Survival Times in Women with Breast Cancer 2026-05-06T05:45:38+00:00 Rahmadani rahma0danird@gmail.com Vinna Sulvia vinnasulvialubis@gmail.com Fathina Nafisa fthnanafisaa13@gmail.com Septrina Kiki Arisandi septrinakikiars@gmail.com Tessy Octavia Mukhti tessyoctaviam@fmipa.unp.ac.id <p><em>Breast cancer remains one of the leading causes of cancer-related mortality worldwide, highlighting the importance of identifying factors that influence patient survival time. Variations in clinical outcomes among patients indicate the need for appropriate statistical methods to evaluate prognostic factors. This studi aims to analyze factors affecting the survival time of breast cancer patients using the Cox Propotional Hazard (Cox PH) model. The data consist of breast cancer patient record with several predictor variabel, including age at diagnosis, type of breast surgery, chemotherapy, hormone therapy, Nottingham Prognostic Index, and tumor size. The analysis procedure includes testingthe propotional hazards assumption and assessing parameter significance using the likelihood ratio test for simultaneous affect and the wald test for partial effect. The resuls show that the propotional hazards assumption is satisfied, indicating that the Cox PH model is appropriate for the data. Simultaneous testing reveals that at least one predictor significanly affect survuval time, while partial testing identifies type of surgery, chemotherapy as significant factors. The hazard ratio estimates indicate that patients undergoing mastectomy have a lower risk of death compared to those receiving breast-conserving surgery. Conversely, chemotherapy and hormone theraoy are associated with a higher risk of death, wich may reflect the more severe clinical conditions of patients receiving these treatments. In conclusion, the Cox PH model provides a reliable approach for identifying key factors influetncing breast cancer survival and offers important implications for clinical decision-making and treatment planning.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Rahmadani, Vinna Sulvia, Fathina Nafisa, Septrina Kiki Arisandi, Tessy Octavia Mukhti https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/486 IHSG Closing Price Prediction on the Indonesian Stock Exchange using the Geometric Brownian Motion Model 2026-04-21T08:28:10+00:00 Sukra Hamna sukrahamna07@gmail.com Devni Prima Sari devniprimasari@fmipa.unp.ac.id <p><em>Being among the leading primary benchmarks reflecting the health of the equity market in Indonesia, the Jakarta Composite Index (IHSG) experiences ongoing price movements shaped by a wide spectrum of domestic and international forces. The inherent unpredictability of these movements underscores the critical need for reliable forecasting methods to guide investors in their decision-making process. In response to this, the present study applies the Geometric Brownian Motion model as a tool for projecting the daily closing values of the IHSG, owing to its well-recognized ability to represent the random characteristics inherent in financial time series. The dataset utilized comprises daily closing price records of the IHSG throughout 2025. The analysis includes the calculation of log returns, normality testing using the Kolmogorov-Smirnov test, and estimation of drift and volatility parameters. Forecasting is performed using simulation with 50 and 1000 iterations, where the initial value is based on the last observed closing price. The findings reveal that the GBM model demonstrates a solid capacity to reflect the volatile behavior of IHSG movements, yielding MAPE figures of 4.50% and 2.81%, which correspond to a very high level of predictive precision. A greater number of iterations was found to produce more consistent and dependable projections, while the estimated values broadly align with the overall trajectory of historical data, notwithstanding the element of randomness embedded in the model. Therefore, the GBM model can be considered an effective method for forecasting stock price movements, particularly for highly volatile market indices such as the IHSG.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Sukra Hamna, Devni Prima Sari https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/491 Mapping Anxiety, Developing Solutions: A Statistical Study of Student Anxiety Using The K-Modes Clustering Method 2026-04-27T03:28:34+00:00 Fadhilah Fitri fadhilahfitri14@gmail.com Fitri Mudia Sari fitrimudiasari@fmipa.unp.ac.id Fauziah Taslim fauziahtaslim@fpk.unp.ac.id Sri Wahyuni 513sriwahyuni@gmail.com <div><em><span lang="IN">Statistics anxiety is a common issue among university students that can negatively affect their learning process and academic performance. This study aims to identify patterns of statistics anxiety among undergraduate students at Universitas Negeri Padang using the Statistics Anxiety Rating Scale (STARS), which consists of six dimensions.</span><span lang="EN-US"> A total of 479 valid responses were analyzed using the k-modes clustering method, which is appropriate for categorical data. The optimal number of clusters was determined using the elbow and silhouette methods, resulting in three clusters. The clustering results reveal three distinct groups of students characterized by high, moderate, and low levels of statistics anxiety. The average silhouette value of 0.52 indicates a moderately well-defined cluster structure. Further analysis shows that each cluster exhibits different patterns across the six anxiety dimensions, highlighting the heterogeneity of students’ responses to statistics. These findings suggest that clustering provides a more informative approach than conventional descriptive analysis in understanding statistics anxiety. The results of this study can serve as a basis for developing targeted strategies to reduce student anxiety in statistics learning</span></em></div> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Fadhilah Fitri, Fitri Mudia Sari, Fauziah Taslim, Sri Wahyuni https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/492 Evaluating Local Parameter Reliability in Hierarchical Geographically Weighted Regression: A Bootstrap and Sign Consistency Approach 2026-05-28T23:13:59+00:00 Fitri Mudia Sari fitrimudiasari@fmipa.unp.ac.id Muhammad Nur Aidi muhammadai@apps.ipb.ac.id Agus Mohamad Soleh agusms@apps.ipb.ac.id Farit Mochamad Afendi fmafendi@apps.ipb.ac.id <p><em>The Hierarchical Geographically Weighted Regression (HGWR) model is widely used to capture spatial heterogeneity and hierarchical data structures simultaneously. However, the reliability of its local parameter estimates remains a critical issue due to potential variability across locations. This study aims to evaluate the reliability of local parameters in the HGWR model using a bootstrap-based approach combined with sign consistency analysis. A cluster bootstrap procedure at the higher-level grouping structure was implemented to generate empirical distributions of parameter estimates, enabling the assessment of statistical significance through confidence intervals. In addition, sign consistency was employed to examine the stability of the direction of local effects across bootstrap replications. The results show that while some local parameters are statistically significant, they do not always exhibit consistent directional effects, indicating potential instability. Conversely, several parameters demonstrate both statistical significance and high sign consistency, suggesting robust local relationships. These findings highlight that relying solely on statistical significance may lead to misleading interpretations of local effects in HGWR models. The combination of bootstrap and sign consistency provides a more comprehensive framework for assessing parameter reliability. This approach contributes to improving the interpretability and robustness of spatial multilevel modeling, particularly in applications involving complex hierarchical and spatial data.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Fitri Mudia Sari, Muhammad Nur Aidi, Agus Mohamad Soleh, Farit Mochamad Afendi https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/487 A Self-Organizing Map Approach for Clustering Provinces Based on Multisectoral Indicators of Stunting Determinants 2026-05-26T03:54:03+00:00 Admi Salma admisalma1@fmipa.unp.ac.id Riwi Dyah Pangesti riwi.pangesti@unib.ac.id Reny Wulandari renywulandari58@gmail.com <p><em>Stunting remains a national as well as a global issue in Indonesia.&nbsp; It becomes one of the key priorities outlined in the Sustainable Development Goals (SDGs). The heterogeneity of multisectoral conditions across provinces also contributes to the variation in stunting prevalence in Indonesia. The implementation of uniform policies to address stunting may not yield optimal results due to the diverse needs of each province. Therefore, </em><em>specific interventions are required to overcome stunting issues. Based on this condition, it is important to cluster provinces based on their characteristics so that the government can determine appropriate interventions for each provincial cluster. Visualization of stunting conditions and multisectoral indicators can also enrich the understanding of each cluster. This study aims to construct clusters of provinces with similar characteristics in terms of multisectoral indicators of stunting determinants. This study applies cluster analysis using a Self-Organizing Map (SOM) algorithm to group provinces. The research steps include data preprocessing, clustering using the SOM algorithm, SOM mapping, and cluster characterization analysis. The results of this study show that three clusters were obtained. The first cluster consists of three provinces characterized by a high maternal mortality rate and a high percentage of exclusive breastfeeding. The second cluster includes nine provinces and is characterized by high risks in maternal and child health as well as economic vulnerability. In addition, the third cluster consists of 26 provinces characterized by relatively good living conditions and quality education.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Admi Salma, Riwi Dyah Pangesti, Reny Wulandari https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/493 Poverty Modeling in East Nusa Tenggara Using Fourier Nonparametric Regression with Cosine–Sine Comparison and Hypothesis Testing 2026-05-18T02:22:12+00:00 Narita Yuri Adrianingsih Naritayuria98@gmail.com Andrea Tri Rian Dani andreatririandani@fmipa.unmul.ac.id I Nyoman Budiantara nyomanbudiantara65@gmail.com Vita Ratnasari vitaratnasari.its@gmail.com Yossy Candra yondra2022@gmail.com Bintang A. Banewang bintang.aries1104@gmail.com Leti S. Gaimau lethygaimau22@gmail.com <p>Kemiskinan adalah isu multidimensional yang kompleks dan tetap menjadi tantangan pembangunan utama di Indonesia, khususnya di Nusa Tenggara Timur (NTT), yang secara konsisten mencatat salah satu tingkat kemiskinan tertinggi secara nasional. Pendekatan parametrik konvensional, seperti regresi linier, seringkali tidak memadai untuk menangkap hubungan nonlinier dan kompleks antara faktor sosioekonomi dan tingkat kemiskinan. Oleh karena itu, penelitian ini mengusulkan pendekatan regresi nonparametrik berdasarkan deret Fourier untuk memodelkan kemiskinan di NTT. Kebaruan penelitian ini terletak pada perbandingan sistematis antara komponen Fourier berbasis kosinus dan berbasis sinus dalam kerangka regresi nonparametrik, dikombinasikan dengan pengujian statistik inferensial untuk mengidentifikasi determinan kemiskinan yang signifikan. Penelitian ini menggunakan data cross-sectional dari 22 kabupaten/kota di NTT untuk tahun 2025. Estimasi model dilakukan menggunakan metode Ordinary Least Squares (OLS), sedangkan parameter osilasi optimal ditentukan menggunakan Generalized Cross-Validation (GCV). Kinerja model dievaluasi menggunakan MSE, RMSE, MAPE, dan koefisien determinasi (R²). Hasil penelitian menunjukkan bahwa model Fourier berbasis kosinus dengan tiga osilasi mengungguli model berbasis sinus, mencapai MSE sebesar 1,903, RMSE sebesar 1,379, MAPE sebesar 5,817%, dan R² sebesar 95,146%. Pengujian hipotesis menunjukkan bahwa semua variabel prediktor secara signifikan memengaruhi tingkat kemiskinan baik secara simultan maupun parsial. Temuan ini menunjukkan bahwa pendekatan regresi nonparametrik Fourier sangat efektif dalam menangkap pola kemiskinan yang kompleks dan berfluktuasi, serta memberikan model yang lebih akurat dan mudah diinterpretasikan untuk mendukung kebijakan pengentasan kemiskinan yang tepat sasaran.</p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Narita Yuri Adrianingsih, Andrea Tri Rian Dani, I Nyoman Budiantara, Vita Ratnasari, Yossy Candra, Bintang A. Banewang, Leti S. Gaimau https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/494 Agricultural Involution in Indonesia: A Generalized Structured Component Analysis (GSCA) Approach with Land and Labor Interaction Effects 2026-05-18T01:47:18+00:00 Urwawuska Ladini urwawuskaladini@uinjambi.ac.id Sella Nofriska Sudrimo sellans@iainsorong.ac.id Dahlia Misrika dahlia.misrika@sci.unand.ac.id <p><em>Indonesian agriculture faces an economic paradox where the sector remains a primary employer despite low wages and stagnant GDP contributions compared to industry. This study analyzes the phenomenon of agricultural involution in Indonesia from 2017 to 2023, specifically examining the interaction effects between land and labor. Utilizing Generalized Structured Component Analysis (GSCA) with an Alternating Least Squares (ALS) approach, the research models the structural relationships between land capacity, labor intensity, productivity, and aggregate output across 34 provinces. The results indicate that land capacity is the dominant determinant of agricultural output with a path coefficient of 0.958, signaling that growth remains extensive rather than intensive. Crucially, labor intensity is found to have a significant negative effect on productivity and total output, confirming the law of diminishing marginal returns and the presence of labor surpluses that exceed optimal points. Furthermore, the interaction between land and labor yields a significant negative coefficient (-0.109), proving that demographic pressure on limited land exacerbates inefficiency and output destruction. Spatial post-hoc analysis indicates that agricultural involution is no longer confined to Java but has evolved into a national phenomenon, as demonstrated by the absence of significant disparities in labor-to-land ratios and productivity between Java and other regions. These findings suggest that sustainable transformation requires integrated policies for land protection, labor restructuring toward non-agricultural sectors, and technological modernization to break the cycle of involution.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Urwawuska Ladini, Sella Nofriska Sudrimo, Dahlia Misrika https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/495 Comparison of District/City Clusters in West Sumatra Province 2019–2025 Based on Labor Indicators Using K-Means Method 2026-05-29T10:19:15+00:00 Naila Marettania nailatania03@gmail.com Zilrahmi zilrahmi@fmipa.unp.ac.id Mellisa Ayuningtyas mellisa@bps.go.id <p style="text-align: justify;"><em><span lang="EN-US" style="font-size: 9.0pt;">This study aims to analyze the clustering of districts/cities in West Sumatra Province based on labor market indicators, namely the Open Unemployment Rate (OUR) and the Labor Force Participation Rate (LFPR for the period 2019–2025. The data used were sourced from the Central Statistics Agency and cover 19 districts/cities. The analysis method employed was C-Means clustering, a method of grouping data based on similarity using Euclidean distance. Two clusters were identified to facilitate interpretation and inter-year comparisons. The results of the study indicate that the regencies/cities are divided into two clusters with distinct characteristics: Cluster 1, characterized by relatively favorable labor market conditions (low OUR and high LFPR), and Cluster 2, characterized by the opposite conditions (high OUR and low LFPR). Cluster membership changes from year to year, indicating the dynamic nature of labor market conditions across regions. Validation results using the Silhouette Coefficient yielded a value of 0.54, indicating a reasonably good quality of clustering. It is hoped that these research findings can serve as a basis for formulating more targeted labor policie.</span></em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Naila Marettania, Zilrahmi, Mellisa Ayuningtyas https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/499 Classification of Stroke Desease Using the Learning Vector Quantization Algorithm 2026-05-25T03:42:42+00:00 Andriarmi Andriarmi andri.armi03@gmail.com Chairina Wirdiastuti chairinawirdiastuti01@gmail.com Syafriandi Syafriandi syafriandi_math@fmipa.unp.ac.id <p><em>Stroke is one of the leading causes of death and disability worldwide, making early detection crutial for timely and appropriate medical treatment. In clinical practice, stroke diagnosis is generally carried out through medical examinations and patient history analysis, but this process is time-consuming and depends on the subjective judgment of medical personnel. Therefore, machine learning approaches can be utilized to support disease classification more quickly and objectively. This study aims to analyze the performance of the Learning Vector Quantization (LVQ) method in classifying stroke disease using a dataset obtained from Kaggle. The dataset used in this study is imbalnced;therefore, the SMOTE (Synthetic Minority Over-Sampling Technique) method was applied to handle class imbalance. The research stages included data preprocessing, splitting data into training and testing sets, LVQ model training, parameter optimization using learning rate and maximum epoch, and model evaluation using accuracy and sensitivity. The result show that the LVQ model trained on the original dataset achived an accuracy of 95,72%, but failed to detect stroke cases with a sensitivity of 0%. After applying SMOTE, the best model achived a stroke sensitivity of 90%, although the accuracy decreased to 49,49% due to the high number of false positives. These findings indicate that LVQ is highly sensitive to data distribution and model parameters, making its performance on this dataset less optimal for stroke classification and more suitable as an initial screening tool.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Andriarmi Andriarmi, Chairina Wirdiastuti, Syafriandi Syafriandi https://ujsds.ppj.unp.ac.id/index.php/ujsds/article/view/496 K-Medoids Clustering Analysis of Regional Development in West Sumatra Based on Socioeconomic Indicators 2026-05-30T14:35:48+00:00 Kayla Faradina kaylafaradina0@gmail.com Fadhilah Fitri fadhilahfitri@fmipa.unp.ac.id <p><em>Regional development disparities among districts and cities in West Sumatra Province remain a persistent challenge, reflected in significant differences across economic, social, and employment indicators. This study aims to cluster 19 districts/cities in West Sumatra Province based on socioeconomic indicators using the K-Medoids clustering method. The variables include GRDP per capita, economic growth rate, GRDP percentage distribution, Human Development Index (HDI), poverty rate, and open unemployment rate, using 2024 data obtained from the Central Bureau of Statistics (BPS) of West Sumatra Province. The optimal number of clusters was determined using the Elbow method, resulting in three clusters. Cluster 1 consists of 12 districts characterized by the lowest average GRDP per capita and HDI, along with the highest poverty rate. Cluster 2 comprises only Kota Padang, which recorded the highest values across most indicators including GRDP per capita, economic growth rate, and HDI, yet also exhibited the highest open unemployment rate. Cluster 3 includes 6 cities with relatively high HDI and the lowest poverty rate among the three clusters. Cluster validation using the Davies-Bouldin Index (DBI) produced a value of 0.8341, indicating that the clustering results are optimal. The findings are expected to provide a reference for local governments and the Regional Development Planning Agency (Bappeda) of West Sumatra Province in formulating more targeted regional development policies based on the characteristics of each cluster.</em></p> 2026-05-31T00:00:00+00:00 Hak Cipta (c) 2026 Kayla Faradina, Fadhilah Fitri