Day 14 (7/25/19): Running the Experiments

The models are still training today although some results can be seen in the early training batches. As expected, catastrophic forgetting is occurring very noticeably in the offline model which is resulting in a drastic decrease in OOD performance as well. I am excited to see how the other models with more incremental learning capabilities turn out with regards to OOD performance in the future.

This graph is showing the accuracy of the model on the training data through 3 batches. Notice how with each new batch of classes, performance significantly drops off as the model struggles to fit to the new data.

Tomorrow I hope to continue the experiments to include the Mahalanobis (distance in feature space) inference method in addition to the Softmax Thresholding method.

Comments

Day 30 (8/16/19): Final Presentations

Today we gave our final presentations, and everyone did a great job. I would like to thank everyone who helped me with this amazing experience! I'm very thankful to have had the opportunity to work on such interesting research with such amazing people this summer.

Day 22 (8/6/19): Streaming Linear Discriminant Analysis

Today I tested the previously trained models using the Stanford dogs dataset as the inter dataset evaluator for OOD instead of the Oxford flowers dataset. However, as expected, the omega values for performance were pretty much the same as before and didn't make much of a difference as the datasets varied. I also implemented a streaming linear discriminant analysis model (SLDA) which differed from the previous incrementally trained models. This model didn't perform as well in terms of accuracy however as only the last layer of the model was trained and streaming is more of a difficult task. Nevertheless, we did show that Mahalanobis can be used in a streaming paradigm to recover some OOD performance in an online setting. This is likely to be a large focus of my presentation as it has never been discussed prior. Tomorrow, I plan to implement an L2SP model with elastic weight consolidation as well as iCarl to serve as two more baselines to compare our experiments to.

Day 24 (8/8/19): Multilayer Perceptron Experiment

I continued gathering more results for my presentation today, and the data table is coming along nicely. We are able to see a significant trend that using Mahalanobis instead of Baseline Thresholding recovers much of the OOD recognition that is lost with streaming or incremental models. The SLDA model appears to be a lightweight, accurate streaming model which can be paired with Mahalanobis to be useful as an embedded agent in the real world. For the purposes of demonstrating catastrophic forgetting, I ran five experiments and averaged the results for a simple incrementally trained MLP. Obviously, the model failed miserably and was achieving only about 1% of the accuracy of the offline model. Including this is only to show how other forms of streaming and incremental models are necessary to develop lifelong learning agents. A diagram of a simple multilayer perceptron.

RIT CIS Summer Internship

Search This Blog