dataset | Academic Journals and Conferences

Method for Detection of Disinformation Based on Text Data Analysis Using TF-IDF and Contextual Vector Representations

The article considers an approach to detecting fake news in the digital environment through text analysis using machine learning and natural language processing methods. The proposed method is based on a hybrid text representation combining frequency features (TF-IDF) and contextual embeddings obtained using the IBM Granite model. A complete data processing cycle was developed, covering the stages of exploratory analysis (EDA), text preprocessing and tokenization, forming vector representations, training a logistic regression model, and obtaining key metrics.

Research of data mining methods for classification of imbalanced data sets

With the rapid development of information technology, which is widely used in all spheres of human life and activity, extremely large amounts of data have been accumulated today. By applying machine learning methods to this data, new practically useful knowledge can be obtained. The main goal of this paper is to study different machine learning methods for solving the classification problem and compare their efficiency and accuracy.

Speech Models Training Technologies Comparison Using Word Error Rate

The main purpose of this work is to analyze and compare several technologies used for training speech models, including traditional approaches as Hidden Markov Models (HMMs) and more recent methods as Deep Neural Networks (DNNs). The technologies have been explained and compared using word error rate metric based on the input of 1000 words by a user with 15 decibel background noise. Word error rate metric has been ex- plained and calculated. Potential replacements for com- pared technologies have been provided, including: Atten- tion-based, Generative, Sparse and Quantum-inspired models.