Published records

Monta Vista Research Club Journal

Demo showcase

Published student research, alongside invented projects that demonstrate the platform.

Papers, 2023 5

Social science · Paper · July 2023

Inferring Hate Speech Trends for Contemporary Tweets Using a Novel Machine Learning Approach from Supervised Learning Algorithms

Aryan Singhal

Social media platforms such as Twitter have become ubiquitous in our contemporary society, providing a platform for individuals to express their opinions and engage in discussions on a wide range of topics including those that are neutral and controversial. However, the growing popularity of Twitter has also led to an increase in the prevalence of hate speech, which raises concerns about its impact on individuals and society. This research investigates hate speech trends on Twitter by utilizing supervised machine learning classification, specifically Naive Bayes, and employing Natural Language Processing (NLP) features such as on a range of neutral and controversial topics. The study compares the prevalence of hate speech in these topics and tracks such trends from January 2022 to January 2023. The results show that hate speech was nearly 400% more prevalent in controversial topics than in neutral topics over the course of the year. In addition, this research finds that controversial topics are consistently more vulnerable to hate speech throughout the course of the year when compared to neutral topics. To conduct this study, a Multinomial Naive Bayes classification model was trained on a publicly available Twitter dataset that was specifically labeled for semantic hate speech and achieved an accuracy rate of 94.46%. Ultimately, the higher vulnerability of controversial topics should necessitate policymakers to introduce stricter warnings or frequent policy reminders to platform users. Such changes will foster a respectful and inclusive platform for users, preserving their freedom of expression and encouraging constructive discussions.

Neuroscience · Paper · June 2023

Detection of Parkinson's disease using Breathing Signals

Advaith Anand

Parkinson's disease (PD) is a progressive neurodegenerative disorder that affects millions of people worldwide. The disease primarily impacts the dopaminergic neurons in the substantia nigra, leading to motor symptoms such as tremors, muscle rigidity, and loss of balance. Early diagnosis of PD is crucial for better management of the disease, allowing for early initiation of treatment and improved patient outcomes. Currently, there is no definitive test for PD; diagnosis relies primarily on the evaluation of clinical symptoms, which often appear several years after the onset of the disease. This delayed detection limits the potential for early intervention and disease management. As a result, there is a growing need for non-invasive, cost-effective, and reliable methods to detect PD at an early stage. In this study, I propose a novel approach for early PD detection by capturing respiratory breathing patterns during sleep using smartphone-generated ultrasonic rays. I employ a Deep Neural Network (DNN) model combining a convolutional neural network (CNN) and a recurrent neural network (RNN) for the classification of the breathing signals into two classes: PD and control.

Neuroscience · Paper · June 2023

Deep Learning Pose Estimation Model for Parkinsonism and Levodopa-Induced Dyskinesia

Yashnil Saha

Diagnosing Parkinson's disease is one of the largest challenges healthcare systems face due to the absence of a specific test for the condition and symptoms varying widely from person to person. Designing an automated model to aid in early diagnosis would greatly contribute to solving this problem. Currently, diagnosis for PD relies on clinical evaluation which has an error rate of approximately 20%, indicating the urgent need for an automated system to be developed. Levodopa is used for the treatment of Parkinson's Disease (PD) but can lead to motor complications known as levodopa-induced dyskinesia (LID) when taken for too long. PD and LID are evaluated according to the Unified Parkinson's Disease Rating Scale (UPDRS) and Unified Dyskinesia Rating Scale (UDysRS) scales, respectively, which range from 0 to 4 (0-normal, 4-severely impaired). The tests are conducted by medical personnel and are very subjective. The goal of this project was to design an algorithm using deep learning for assessment of parkinsonism and LID using pose estimation. Two models were created: a regression model to predict the clinical rating from 0 to 4 and a classification model to determine whether the patient had PD or LID. During the feature extraction process, 32 features were extracted per joint trajectory including 15 kinematic, 16 spectral, and the convex hull of the movements. Then, the two neural network models were trained on these features to be able to predict their respective targets. The classification model achieved a mean F1-score greater than 0.8 and the regression model attained a root mean square error less than 0.550, proving that this project was a promising start in the venture to automate diagnosis of Parkinson's disease.

Engineering and robotics · Paper · June 2023

Pillar-Based Overhang Generation to Reduce Waste in 3D Printing

Raymond Feng

Pillar-Based Overhang Generation, or PBOG, is a novel algorithm that generates support structures for 3D printing. In 3D printing, overhangs are parts in midair that cannot be printed. The most popular way to overcome this issue is to add support structure. The PBOG algorithm is proposed to decrease material usage in support structure. The algorithm uses three steps: overhang detection, vertex simplification, and pillar generation. Compared with Cura and Meshmixer, results show that PBOG uses less waste and has a higher success rate. In conclusion, PBOG outperforms currently popular methods Cura and Meshmixer in material waste and success rate.

Astronomy and astrophysics · Paper · June 2023

Correcting Mislabeled Quasars in Extragalactic Catalogs

Arjun Shrivastava

Quasars, a type of active galactic nuclei (AGNs), are some of the brightest objects in the universe. They allow astronomers to accurately observe distant objects and look farther back in time, offering researchers a better understanding of our universe's history. However, the X-ray and extragalactic databases that catalog astronomical objects often mislabel quasars as other objects. Therefore, I seek to improve these classifications by identifying quasars in the Deep Fields component of the Canada-France-Hawaii Telescope Legacy Survey (CFHTLS) database and cross-checking them with the extragalactic catalogs. The typical method of classifying an object as a quasar is via manual visual classification: examining unfolded light curves for star-like objects that exhibit irregular short-term and long-term variations in brightness. Unfortunately, this is time-intensive and prone to human error, so I attempted to accelerate the process and boost accuracy with deep neural networks. I first visually classified data of about 4,000 objects for my training set and included four different filters. After preprocessing, I trained a Long Short-Term Memory (LSTM) neural network with different variations of hyperparameters until I achieved an accuracy of 97.5%. When I ran my model through my data, it identified 992 new quasars, 796 of them being actually quasars while the rest were misclassified, yielding an overall binary accuracy of 80% for the entire CHFTLS Deep Field database. Of the identified quasars, 14.8% of which were new quasars and 83.3% were mislabeled in the NASA/IPAC Extragalactic Database (NED). Many of the mislabeled quasars tended to appear as galaxies as well as unidentified sources of ultraviolet or X-ray radiation. In the future, I seek to improve my model for higher accuracy.