Computer science · Paper · August 2021
Jai Sharma, Milind Maiti, Christopher Sun
Dropout Regularization, serving to reduce variance, is nearly ubiquitous in Deep Learning models. We explore the relationship between the dropout rate and model complexity by training 2,000 neural networks configured with random combinations of the dropout rate and the number of hidden units in each dense layer, on each of the three data sets we selected. The generated figures, with binary cross entropy loss and binary accuracy on the z-axis, question the common assumption that adding depth to a dense layer while increasing the dropout rate will certainly enhance performance. We also discover a complex correlation between the two hyperparameters that we proceed to quantify by building additional machine learning and Deep Learning models which predict the optimal dropout rate given some hidden units in each dense layer. Linear regression and polynomial logistic regression require the use of arbitrary thresholds to select the cost data points included in the regression and to assign the cost data points a binary classification, respectively. These machine learning models have mediocre performance because their naive nature prevented the modeling of complex decision boundaries. Turning to Deep Learning models, we build neural networks that predict the optimal dropout rate given the number of hidden units in each dense layer, the desired cost, and the desired accuracy of the model. Though, this attempt encounters a mathematical error that can be attributed to the failure of the vertical line test. The ultimate Deep Learning model is a neural network whose decision boundary represents the 2,000 previously generated data points. This final model leads us to devise a promising method for tuning hyperparameters to minimize computational expense yet maximize performance. The strategy can be applied to any model hyperparameters, with the prospect of more efficient tuning in industrial models.
Astronomy and astrophysics · Paper · June 2021
Iona Xia
The universe is still largely a mystery to scientists. Quasars are one such mystery, whose emission spectra produce absorption lines such as Ca II when passing through gas of distant galaxies. This data helps astronomers understand more about interstellar gas, dust, and galaxy and star formation and evolution (including my Milky Way). However, these current absorber databases are extremely limited, and traditional methods make them hard to detect. Thus, I seek to discover more Ca II absorbers by developing deep neural networks, which are more accurate and faster. I first found absorbers traditionally to produce a test set. I cropped, normalized, and handpicked through thousands of spectra and discovered 256 Ca II original absorbers. To obtain large training sets, I generated tens of thousands of artificial samples by inserting Ca II lines at corresponding wavelengths in real spectra. I preprocessed the data and created neural network models after testing different hyperparameter configurations. Overall, my accuracy for absorber detection is 95% (Ca II), 15 times higher than traditional methods, and I added significant amounts of new absorbers to the current dataset for Ca II, completing my goal. As for challenges, I concluded that most false negatives are due to noise and weak lines. Furthermore, my discovered absorbers agree with statistical tests of previous studies. In the future, I plan to discover more absorbers using my models and run statistical studies on them.
Chemistry and materials · Paper · June 2021
Sanjana S. Jilla
One of the biggest threats our planet faces is the threat of global warming. Electrical power has become crucial to our modern lifestyle, but today, electricity is generated by burning fossil fuels and coal, which has many harms and disadvantages associated with it. Fortunately, there are multiple alternative energy sources, including sunlight, a powerful, inexhaustible, and clean resource. There are several different types of solar cells, but one of the most efficient and low-cost types is the perovskite solar cell. Perovskite solar cells (PSCs) have recently received considerable attention due to the high energy conversion efficiency achieved within a few years of their inception. However, today, the most common perovskite is methyl ammonium iodide (MAPbI3), which contains levels of toxic lead. The science community has been searching for lower-toxicity perovskite-type materials, but testing all of the possible lead-free perovskites requires a huge amount of time and funding. Recent advances in computing power have enabled the generation of large datasets for materials and data-driven approaches to problem-solving in materials science, including materials discovery. Machine learning is the primary tool for manipulating such large datasets, predicting unknown material properties and uncovering relationships between structure and property. The goal of this project is to create a Machine Learning (ML) driven software system that increases the efficiency of the solar cell design process. I will do this by identifying the best perovskites by optimizing material composition and determining the importance of the features of each element in the perovskite to the overall efficiency of the PSC. The Machine Learning (ML) driven software system needs to accurately predict key information such as the heat of formation (delta Hf) and band gap (Eg) and accurately use the training data to form an accurate prediction of the best materials to form a perovskite solar cell (Im). The models must predict with 90% accuracy of prediction for the project to be successful. This program is written in Python, and uses Machine Learning to make predictions about the heat of formation and bandgap of various double halide perovskites.
Biology and biomedicine · Paper · May 2021
Anika Nagavara
According to the Centers for Disease Control and Prevention (CDC), 1 in 13 people have asthma. Each day, ten Americans die from asthma (Asthma and Allergy Foundation of America 2018). A previous project that was conducted last year looked at the correlation between zinc deficiency and asthma and yielded positive results meaning that a correlation between additional zinc intake and a better control of asthma could be seen. This project observed whether there is a correlation between iodine deficiency and asthma since iodine has previously been used to treat inflammatory diseases since it has properties that can stabilize thyroid hormone levels as well as reduce bronchial secretions and mediate immune cell responses (Lake 2017). In order to supply the iodine, potassium iodide was used since iodine is more easily absorbed by the body when it is in the form of potassium iodide. A combination treatment of potassium iodide and zinc was also given.
Biology and biomedicine · Paper · March 2021
Rishi Pankhaniya
Tunicamycin is a commonly used drug to cause an unfolded protein response in multiple myeloma cells in order to treat the cancer. However, many multiple myeloma cell lines have slowly developed resistance to this treatment. The goal of my project is to find out the reasons in the RNA behind why this resistance is caused, in order to create a better, altered treatment that could possibly circumvent these problems. First, multiple myeloma cells were treated with tunicamycin repeatedly four times, such that the living cells would sufficiently have developed resistance to the treatment. Then a short-term treatment was performed before plating the cells, dividing the cells into a control and treated group to find differences between their RNA to find indicators that cause the resistance. After plating the cells and isolating the RNA, a mass transcriptome analysis returned exonic data to allow us to look at how the resistance was being developed through the creation of proteins and certain protein responses. Upon looking at the data, several markers were made clear, such as the suppression of VAPA and DDIT3 as examples. Through looking at all of these gene markers, the treatment can be slightly altered in order to prevent the suppression of certain responses that would cause the cancer cell to die, thus making the treatment apply to a wider range of cells.