Skip to main content

Posts

Day 29 (8/15/19): Final Day Before Presentations

Most of today was also spent practicing and editing my presentation to make it as professional as I can. I'm really looking forward to the opportunity to present my work to faculty and friends tomorrow. Here is a link to the slides for my final presentation: Novelty Detection in Streaming Learning using Neural Networks

Day 28 (7/14/19): Presentation Dry Run

In the morning, all of us interns got the chance to practice our presentations in front of each other in the auditorium. I was pretty happy with how mine went overall but the experience was definitely valuable in identifying typos or slight adjustments that should be made. Throughout the rest of the day, I tried to implement these changes and clean up a few plots that I want to include for Friday.

Day 27 (8/13/19): Improving Presentation Plots

Today I practiced my presentation more and also added better visual graphs to better understand my results. Now, the line graphs show the results after each batch of training so you can see the trend in accuracy and OOD detection over time. Lastly, I added a bar chart at the end of the presentation to summarize my overall results in addition to the spider chart.

Day 26 (8/12/19): Presentation Revisions

Today was very useful for making revisions and edits to my presentation. I ran through it in front of my lab this morning and got lots of helpful feedback as to how to make it more accessible to a general audience (eliminating jargon). Every day I am becoming more and more confident with the talk, and I'm looking forward to presenting on Friday! Furthermore, I learned today that I will be able to get my RIT computer account/email to stay active for a few months after the internship ends. This will allow me to continue communicating with the lab via Slack and help review and write a research paper including some of the work I have pursued over the past six weeks. We hope to submit this paper to get it published at a conference in the fall (possible AAAI).

Day 25 (8/9/19): Finishing Presentation

Today I made a lot of progress finishing up my presentation. I feel like we developed an interesting story to tell around the data we collected from the experiments, and I am excited to get a chance to share my results. Much of the beginning of the presentation is spent explaining high level concepts such as deep learning and machine learning so I will have a better idea of what I will need to include after my meeting with Joe and Amy on Monday. I will continue to keep practicing my presentation over the weekend and possible include more results from iCaRL and MLP w/ EWC models if I can get them trained. Below I have included a visualization of one of the most important results from my project. Notice how the SLDA w/ Mahalanobis model outperforms the other models in accuracy and OOD recognition combined (the more area a model has in the spider plot, the better it performed overall).

Day 24 (8/8/19): Multilayer Perceptron Experiment

I continued gathering more results for my presentation today, and the data table is coming along nicely. We are able to see a significant trend that using Mahalanobis instead of Baseline Thresholding recovers much of the OOD recognition that is lost with streaming or incremental models. The SLDA model appears to be a lightweight, accurate streaming model which can be paired with Mahalanobis to be useful as an embedded agent in the real world. For the purposes of demonstrating catastrophic forgetting, I ran five experiments and averaged the results for a simple incrementally trained MLP. Obviously, the model failed miserably and was achieving only about 1% of the accuracy of the offline model. Including this is only to show how other forms of streaming and incremental models are necessary to develop lifelong learning agents. A diagram of a simple multilayer perceptron.

Day 23 (8/7/19): Averaging SLDA Results

Today I worked a lot on my presentation and ran a few more experiments to include. So far we have averaged data for offline, full rehearsal, and SLDA models. We were hoping to test the EWC model today but didn't get the chance due to some bugs in our code. Hopefully, that will be ready for my presentation next week. Here is a preview of the title of the presentation:

Day 22 (8/6/19): Streaming Linear Discriminant Analysis

Today I tested the previously trained models using the Stanford dogs dataset as the inter dataset evaluator for OOD instead of the Oxford flowers dataset. However, as expected, the omega values for performance were pretty much the same as before and didn't make much of a difference as the datasets varied.  I also implemented a streaming linear discriminant analysis model (SLDA) which differed from the previous incrementally trained models. This model didn't perform as well in terms of accuracy however as only the last layer of the model was trained and streaming is more of a difficult task. Nevertheless, we did show that Mahalanobis can be used in a streaming paradigm to recover some OOD performance in an online setting. This is likely to be a large focus of my presentation as it has never been discussed prior. Tomorrow, I plan to implement an L2SP model with elastic weight consolidation as well as iCarl to serve as two more baselines to compare our experiments to.

Day 21 (8/5/19): Averaging Experiments

This morning I was at a basketball camp so I came into the lab around noon. Much of the day was spent waiting for some models to finish training to I worked on adding some slides to my presentation document. In the afternoon, I got back results from models that I could average together to get more accurate results. The general trend remained the same however which indicated that the Mahalanobis Intra Dataset OOD actually performed better when it was trained incrementally (albeit with full rehearsal) than when it was trained offline. I am not sure yet what the reason for this is, but I will continue to look into it. The green line denotes the intra dataset mahalanobis OOD omega for full rehearsal. Note how it consistently is above 1 even as more batches are learned incrementally.

Day 20 (9/2/19): Fixing OOD Evaluation with Equal Sample Distribution

Today I reran the experiments from yesterday with the dataloaders for the OOD performance evaluation having equal in- and out-loader sample sizes. Theoretically, this would lead to a more accurate AUROC metric. However, just glancing at a visualization of the new results, it appears that we are achieving the same interesting results as yesterday. Unsure of the underlying reason why, I hope to plot the metrics we calculated on the same set of axis to get a better representation of the results.

Day 19 (8/1/19): Analyzing the Results of Early Experiments

Today I reviewed the results of the earlier experiments I ran with my mentor and other students in the lab. The most interesting result (which we most likely will have to repeat to ensure accuracy) was that the intra dataset OOD performance for the full rehearsal model was actually higher than that of the offline model. The y-axis represents the omega value for intra datset OOD with mahalanobis. The x-axis represents the number of classes learned. Today was also the RIT Undergraduate Research Symposium which was very fun to attend. Along with a few other interns, I listened to three presentations which talked about political biases affecting article credibility, fingerprinting as a means of cybersecurity defense, and laughter detection and classification using deep learning respectively. Each talk was interesting in its own way, and I enjoyed learning about the other research being performed in a similar field to the one I am working in. Tomorrow I hope to run more experiments...

Day 18 (7/31/19): Obtaining Results from Early Experiments

Today I reviewed some of the first true results for the early rounds of experiments I performed. For the offline model (intended to be used as the baseline for the calculating the omega values of incrementally learned models), the final batch of 20 classes yielded an accuracy of 81.20%, an AUROC for Gaussian Noise of .99, an AUROC for Inter datset OOD of .82, and an AUROC for Intra dataset OOD of .80. It is important to note as well that I switched the learning rate scheduler to be exponentially defined rather than decaying the learning rate by steps once it reaches 2/3 of the batch iterations. The full rehearsal model, as expected, almost performed as well as the offline model achieving an  accuracy of 78.12%, an AUROC for Gaussian Noise OOD Omega of .89, an AUROC for Inter datset OOD Omega of .92, and an AUROC for Intra dataset OOD Omega of .98. It will be interesting to see how these results compare to future models. Most likely, these less memory-intensive models will per...

Day 17 (7/30/19): Adding Bounding Boxes in Experiments

Today was very successful as we finally were able to finish debugging the script we were using for training our experiment models. Previously, I was seeing the loss function diverge to nan during the second batch and the problem was actually that I was forgetting to shuffle the data with each new batch. After cleaning up much of the script, I was able to organize the code to be more general purpose for future experiments. I ran the new code to train an offline so i could generate some baseline numbers to use for omega values for the experiments. However, I am still adjusting a few hyperparameters (namely the patience counter for early stopping to prevent overfitting) to see what will yield the highest offline accuracy with bounding boxes implemented. The full rehearsal model was fairly simple to implement after the offline model trained so I plan to start evaluating Mahalanobis OOD and more complex models tomorrow.

Day 16 (7/29/19): Achieving Better Results by Adjusting Learning Rate

This morning I reviewed the results of the experiments I ran over the weekend. However, the tensorboard graphs showed that after a certain number of batches, every model appeared to stop learning on the training data, essentially becoming stagnant at the end. Below is a graphical representation of those models over the weekend. I realized that the training script I was using didn't explicitly reset the optimizer and learning rate scheduler after each batch as I intended. Therefore, as the learning rate was continuously getting smaller, the model was no longer learning in the further batches. I adjusted my script to reset the parameters of both the optimizer and scheduler after each batch of 20 classes and restarted many of the experiments. I also added some more capabilities to load and save the models by creating dictionaries to track how the accuracy and omega values (against default values of offline model) change as the number of batches and classes increase. I hope to h...

Day 15 (7/26/19): Testing Models with Rehearsal and L2-SP Regularization

Today I continued the experiments from yesterday along with implementing a L2SP Model and Partial Rehearsal with Baseline OOD. So far it seems that the performance of every model (both accuracy and area under ROC curve) significantly drops as the number of classes learned increases. Implementing the more complex models such as SLDA, S-SVM, and L2SP (EWC) as well as more accurate inference methods such as Mahalanobis will be a challenge but also interesting to see how well they perform. The blue line represents the L2SP model and the red line represents the Full Rehearsal model. These have only been trained for around three batches of 20 classes and will continue to learn overnight. However, the performance trend will most likely continue as the accuracy drops with newly added classes. I also finished my presentation outline today which can be found at this link:  RIT Presentation Outline

Day 14 (7/25/19): Running the Experiments

The models are still training today although some results can be seen in the early training batches. As expected, catastrophic forgetting is occurring very noticeably in the offline model which is resulting in a drastic decrease in OOD performance as well. I am excited to see how the other models with more incremental learning capabilities turn out with regards to OOD performance in the future. This graph is showing the accuracy of the model on the training data through 3 batches. Notice how with each new batch of classes, performance significantly drops off as the model struggles to fit to the new data. Tomorrow I hope to continue the experiments to include the Mahalanobis (distance in feature space) inference method in addition to the Softmax Thresholding method.

Day 13 (7/24/19): Offline and Rehearsal Model Experiments

Today I began some experiments that I hope to include in my final project presentation. The main objective as of now is to figure out which incremental learning strategies yield the best out-of-distribution  (OOD) performance. For the experiments I performed today, I trained all layers of the models in batches of 20 classes (10 batches for the 200 species in CUB200 dataset) and evaluated OOD using a baseline softmax thresholding method. The performance metrics I hope to obtain are the Omega alpha (how accurate model is compared to offline model) and Omega OOD (how accurate the model is at novelty detection compared to offline model). *These models are currently still training so I should have the results in the morning. During lunch I went to the seminar which discussed ASL, specifically how it was important here at RIT. I found the talk very interesting and even learned a few signs which might be useful someday. Tomorrow I hope to continue my work on this project and expand...

Day 12 (7/23/19): Bounding Classifiers

Today I experimented with different bounded classifiers for open set recognition. A bounding classifier essentially is a type of model which can detect out-of-distribution (OOD) samples, i.e. when presented with an image of a class it has not been trained on. In the image below, a bounded classifier would identify images found in the less dense regions as unknowns rather than try to fit them to a previously-learned class. I performed my experiment evaluating different bounding classifiers by testing a Resnet50's accuracy in detecting out-of-distribution samples (original dataset is CUB200) either from generated Gaussian noise or the Oxford Flowers dataset. Here are two sample images from those respective dataloaders: The results I achieved were very similar to those shown in this table (the third row is the CUB200 dataset): Tomorrow I hope to finally begin to look into the intersection of incremental learning and open set recognition (having experiment with b...

Day 11 (7/22/19): Partial Rehearsal and Out-Of-Distribution Recognition

This morning I reviewed the results of the different incremental learning models I started training over the weekend. Here are the results for five different methods. *Accuracy is computed as the average of the batch 1 and batch 2 accuracies to represent the overall test set. Omega is the average ratio between the accuracy of the incremental model vs. the offline model (offline accuracy with default hyperparameters: .7330) No Regularization: Batch 1--> Testing 45/45 Accuracy: 0.0241 Batch 2--> Testing 46/46 Accuracy: 0.7702 Accuracy: 0.3972 Omega: 0.5418 L2 Regularization: Testing 45/45 Accuracy: 0.0132 Testing 46/46 Accuracy: 0.7868 Accuracy: 0.4000 Omega: 0.5457 L2SP Regularization: Batch 1--> Testing 47/47 Accuracy: 0.0000 Batch 2--> Testing 45/45 Accuracy: 0.8185 Accuracy: 0.4093 Omega: 0.5583 Pseudo-Rehearsal (w/ random sampling of ten images per batch-1 class): Testing 45/45 Accuracy: 0.4166 Testing 46/46 Accuracy: 0.7988 ...