A Scalable Transformer-Based Framework for Sentiment Analysis and Public Opinion Mining Across Heterogeneous Social Media Platforms

Authors

  • Fateh Bahadur Kunwar Associate Professor & Dean (Academics & IQAC) Department of Computer Science and Engineering
  • Yash Bansal Student, Department of Computer Science and Engineering Greater Noida Institute of Technology, Greater Noida (U.P.), India – 201310
Published 2026-09-23
Section Research Paper
AccessSubscription

Keywords:

Terms—sentiment analysis social media mining public opinion monitoring transformer models RoBERTa sarcasm detection code-mixed NLP distributed inference platform-adaptive fine-tuning opinion mining

Abstract

Social media platforms generate more than 500 million posts every day, creating a constantly changing record of public opinion about political events, public health crises, brands, and social movements. Although this data is extremely valuable, turning raw social media text into clear and useful opinion insights is still difficult because of three main challenges. First, different platforms such as Twitter/X, Reddit, YouTube, and Facebook use different writing styles, vocabulary, post lengths, and noise patterns. Second, social media language often includes code- switching, sarcasm, implicit sentiment, and platform-specific slang, which makes it hard for traditional lexicon- based or single-dataset trained models to understand the real meaning. Third, analyzing millions of posts in real time requires highly scalable computational systems. To address these challenges, this paper introduces the Cross-Platform Sentiment and Opinion Intelligence Network (CPSOIN), a unified transformer-based framework. CPSOIN combines three key components: Platform-Adaptive Fine-Tuning (PAFT) applied to a RoBERTa-large model to handle differences between platforms, a Sarcasm-Aware Contextual Encoding (SACE) module that detects sarcasm using incongruity between sentence meaning and knowledge-based expectations, and a Distributed Streaming Inference Engine (DSIE) built with Apache Kafka and ONNX Runtime. This system can process up to 47,300 posts per second with latency below 80 milliseconds on standard hardware. The framework is evaluated on five datasets: SemEval-2017 Task 4, SentiRaama, Reddit SARC 2.0, Twitter COVID-19 Sentiment, and a new Indian General Election 2024 opinion dataset. Results show that CPSOIN improves weighted F1 scores by 3.1–8.7 percentage points compared to strong baseline models, reduces sarcasm misclassification by 34.2%, and achieves 4.1× higher inference throughput than standard PyTorch-based systems at production-scale data volumes.

References

  1. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding." in Proc. NAACL-HLT, pp. 4171–4186, 2019.
  2. Y. Liu et al., "RoBERTa: A robustly optimized BERT pretraining approach," arXiv:1907.11692, 2019.
  3. D. Q. Nguyen et al., "BERTweet: A pre-trained language model for English tweets," in Proc. EMNLP (Systems), pp. 9–14, 2020.
  4. Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le, "XLNet: Generalized autoregressive pretraining for language understanding," in Adv. NeurIPS, pp. 5753–5763, 2019.
  5. Y. Yin and H. Zheng,"SentiBERT: A transferable transformer-based architecture for compositional sentiment semantics," in Proc. ACL, pp. 3695–3706, 2020.
  6. M. Sun, X. Li, J. Guo, and Q. Liu, "ABSA-BERT: Aspect-based sentiment analysis using BERT with selective attention," IEEE Access, vol. 9, pp. 93847–93857, 2021.
  7. A. Ghosh and T. Veale,"Fracking sarcasm using neural network," in Proc. 7th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pp. 161–169, 2016.
  8. N. Hazarika, S. M. Poria, R. Mihalcea, E. Cambria, and R. Zimmermann," CASCADE: Contextual sarcasm detection in online discussion forums," in Proc. COLING, pp. 1837–1848, 2018.
  9. A. Oprea and W. Magdy," iSarcasm: A dataset of intended sarcasm," in Proc. ACL-IJCNLP, pp. 1279–1289, 2021.
  10. B. Patra, R. Das, A. Das, and R. Prasath, "Sentiment analysis of code-mixed Indian languages: An overview of SAIL_Code-Mixed shared task @ICON-2017," arXiv:1803.06278, 2018.
  11. S. Patwa et al., "SemEval-2020 task 9: Overview of sentiment analysis of code-mixed tweets," in Proc. SemEval, pp. 774–790, 2020.
  12. Z. Chen and B. Qian,"Transfer capsule network for aspect level sentiment classification," in Proc. ACL, pp. 547–556, 2019.

Published

2026-09-23

How to Cite

A Scalable Transformer-Based Framework for Sentiment Analysis and Public Opinion Mining Across Heterogeneous Social Media Platforms. (2026). NOLEGEIN-Journal of Advertising and Brand Management, 9(2). https://mbajournals.in/index.php/JoABM/article/view/2061

Similar Articles

41-50 of 57

You may also start an advanced similarity search for this article.