From 11dbac57e92e8ecfc1f064193228f1ad9afaf303 Mon Sep 17 00:00:00 2001
From: SpeedProg <speedprog@googlemail.com>
Date: Thu, 02 Jan 2020 16:04:02 +0000
Subject: [PATCH] removed weights

---
 README.md |  178 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
 1 files changed, 172 insertions(+), 6 deletions(-)

diff --git a/README.md b/README.md
index 2c38786..55a9a25 100644
--- a/README.md
+++ b/README.md
@@ -1,10 +1,35 @@
+# Magic: The Gathering Card Detector
 
-# Magic: The Gathering Card Detection Model
+MTG Card Detector is a real-time application that can identify Magic: The Gathering playing cards from either an image or a video. It utilizes various computer vision techniques to process the input image, and uses [perceptual hashing](https://jenssegers.com/61/perceptual-image-hashes) to identify the detected image of the cards with the matching cards from the database of MTG cards. Refer to [opencv_dnn.py](https://github.com/hj3yoo/mtg_card_detector/blob/master/opencv_dnn.py) for more detailed implementation.
 
-This is a fork of [Yolo-v3 and Yolo-v2 for Windows and Linux by AlexeyAB](https://github.com/AlexeyAB/darknet#how-to-compile-on-linux) for creating a custom model for [My MTG card detection project](https://github.com/hj3yoo/MTGCardDetector). 
+**Demo:**
+
+[![Demo #1](https://img.youtube.com/vi/BZkRZDyhMRE/0.jpg)](https://www.youtube.com/watch?v=BZkRZDyhMRE "Demo #1")
+
+You can run the demo using the following:
+
+```
+python3 opencv_dnn.py [-i path/to/input/file -o path/to/output/directory -hs (one of 16/32)  -dsp -dbg -gph]
+```
+
+Initially, the project used a powerful neural network named ['You Only Look Once (YOLO)'](https://arxiv.org/pdf/1506.02640v5.pdf) to detect individual cards, but it has been removed as of Oct 12th, 2018 [(note)](https://github.com/hj3yoo/mtg_card_detector#oct-12th-2018) in favour of classical CV techniques. 
+
+**Demo:**
+
+[![Demo #2](https://img.youtube.com/vi/kFE_k-mWo2A/0.jpg)](https://www.youtube.com/watch?v=kFE_k-mWo2A "Demo #2")
+
+You can still find the files used to train them:
+
+- [tiny_yolo.cfg](https://github.com/hj3yoo/mtg_card_detector/blob/master/cfg/tiny_yolo.cfg)
+- [tiny_yolo_final.weights](https://github.com/hj3yoo/mtg_card_detector/blob/master/weights/second_general/tiny_yolo_final.weights)
+- [obj.data](https://github.com/hj3yoo/mtg_card_detector/blob/master/data/obj.data) and [obj.names](https://github.com/hj3yoo/mtg_card_detector/blob/master/data/obj.names)
+- [fetch_data.py](https://github.com/hj3yoo/mtg_card_detector/blob/master/fetch_data.py): aggregates card images and database from [scryfall.com](https://scryfall.com/)
+- [transform_data.py](https://github.com/hj3yoo/mtg_card_detector/blob/master/transform_data.py): generate training images using the aggregated card images and database
+- [setup_train.py](https://github.com/hj3yoo/mtg_card_detector/blob/master/setup_train.py): create train.txt and test.txt required to train YOLO from the training dataset
+
+---------------------------------------------------------------------
 
 ## Day ~0: Sep 6th, 2018
----------------------
 
 Uploading all the progresses on the model training for the last few days.
 
@@ -28,8 +53,7 @@
 
 The second and third problems should easily be solved by further augmenting the dataset with random lighting and image skew. I'll have to think more about the first problem, though.
 
-## Day 1
------------------------
+## Sept 7th, 2018
 
 Added several image augmentation techniques to apply to the training set: noise, dropout, light variation, and glaring:
 
@@ -45,6 +69,148 @@
 
 <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_detection_result_1.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_decision_result_2.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_decision_result_3.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_decision_result_4.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_decision_result_5.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_decision_result_6.jpg" width="360">
 
-<img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_learning_curve.jpg" width="360"> 
+<img src="https://github.com/hj3yoo/darknet/blob/master/figures/1_learning_curve.jpg" width="640"> 
 
 The video demo can be found here: https://www.youtube.com/watch?v=kFE_k-mWo2A&feature=youtu.be
+
+
+## Sept 10th, 2018
+
+I've been training a new model with a full YOLOv3 configuration (previous one used Tiny YOLOv3), and it's been taking a lot more resources:
+
+<img src="https://github.com/hj3yoo/darknet/blob/master/figures/2_learning_curve.jpg" width="640"> 
+
+The author of darknet did mention that full network will take significantly more training effort, so I'll just have to wait. At this rate, it should reach 50k epoch in about a week :/
+
+
+## Sept 13th, 2018
+
+The training for full YOLOv3 model has turned sour - the loss saturated around 0.45, and didn't seem like it would improve in any reasonable amount of time. 
+
+<img src="https://github.com/hj3yoo/darknet/blob/master/figures/3_learning_curve.jpg" width="640"> 
+
+As expected, the performance of the model with 0.45 loss was fairly bad. Not to mention that it's quite slower, too. I've decided to continue with tiny YOLOv3 weights. I tried to train it further, but it was already saturated, and was the best it could get.
+
+---------------------
+
+Bad news, I couldn't find any repo that has python wrapper for darknet to pursue this project further. There is a [python example](https://github.com/AlexeyAB/darknet/blob/master/darknet.py) in the original repo of this fork, but [it doesn't support video input](https://github.com/AlexeyAB/darknet/issues/955). Other darknet repos are in the same situation.
+
+I suppose there is a poor man's alternative - feed individual frames from the video into the detection script for image. I'll have to give it a shot.
+
+
+## Sept 14th, 2018
+
+Thankfully, OpenCV had an implementation for DNN, which supports YOLO as well. They have done quite an amazing job, and the speed isn't too bad, either. I can score about 20~25fps on my tiny YOLO, without using GPU.
+
+
+## Sept 15th, 2018
+
+I tried to do an alternate approach - instead of making model identify cards as annonymous, train the model for EVERY single card. As you may imagine, this isn't sustainable for 10000+ different cards that exists in MTG, but I thought it would be reasonable for classifying 10 different cards.
+
+Result? Suprisingly effective.
+
+<img src="https://github.com/hj3yoo/darknet/blob/master/figures/4_detection_result_1.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/4_detection_result_2.jpg" width="360"><img src="https://github.com/hj3yoo/darknet/blob/master/figures/4_detection_result_3.jpg" width="360"> <img src="https://github.com/hj3yoo/darknet/blob/master/figures/4_detection_result_4.png" width="360">
+
+They're of course slightly worse than annonymous detection and impractical for any large number of cardbase, but it was an interesting approach.
+
+------------------
+
+I've made a quick openCV algorithm to extract cards from the image, and it works decently well:
+
+<img src="https://github.com/hj3yoo/darknet/blob/master/figures/4_detection_result_5.jpg" width="360">
+
+At the moment, it's fairly limited - the entire card must be shown without obstruction nor cropping, otherwise it won't detect at all.
+
+Unfortunately, there is very little use case for my trained network in this algorithm. It's just using contour detection and perceptual hashing to match the card.
+
+
+## Sept 16th, 2018
+
+I've tweaked the openCV algorithm from yesterday and ran for a demo:
+
+https://www.youtube.com/watch?v=BZkRZDyhMRE&feature=youtu.be
+
+## Oct 4th, 2018
+
+With the current model I have, there seems to be little hope - I simply don't have enough knowledge in classical CV technique to separate overlaying cards. Even if I could, perceptual hash will be harder to use if I were to use only a fraction of a card image to classify it.
+
+There is an alternative to venture into instance segmentation with [mask R-CNN](https://arxiv.org/pdf/1703.06870.pdf), at the cost of losing real-time processing speed (and considerably more development time). Maybe worth a shot, although I'd have to nearly start from scratch (other than training data generation).
+
+## Oct 10th, 2018
+
+I've been trying to fiddle with the mask R-CNN using [this repo](https://github.com/matterport/Mask_RCNN)'s implementation, and got to train them with 60 manually labelled image set. The result is not too bad considering such a small dataset was used. However, there was a high FP rate overall (again, probably because of small dataset and the simplistic features of cards).
+
+<img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/5_rcnn_result_1.png" width="360"><img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/5_rcnn_result_2.png" width="360"><img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/5_rcnn_result_3.png" width="360"><img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/5_rcnn_result_4.png" width="360"><img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/5_rcnn_result_5.png" width="360">
+
+Although it may be worth to generate large training dataset and train the model more thoroughly, I'm being short on time, as there are other priorities to do. I may revisit this later. I will be cleaning this repo in the next few days, wrapping it up for now.
+
+## Oct 12th, 2018
+
+I've been able to significantly cut down the processing time of the current implementation. For n cards detected in the video, the latency has decreased from (65+50n)ms to (7+16n)ms. There were two major bottlenecks that was slowing the program down:
+
+--------------------------
+
+In order to identify the card from the snippet of the card image, I'm using perceptual hashing. When the card is detected in YOLO, I compute its pHash value from its image, and compare it with the pHash of every cards in the database to find the match. This process has a speed of O(n * m), where n is the number of cards detected in the image and m is the number of cards in the database. With more than 10000 different cards printed in MTG history, this computation was the first bottleneck. For the 50ms increment per detected card mentioned above, majority of that time was spent trying to subtract two 1024-bit hashes 10000+ times - that's more than 10^10 comparisons right there!
+
+First, there were some overhead that was coming from the implementation of the library. The following is the elapsed time for subtracting pHash for all 10000 elements in pandas database:
+
+| hash_size  | elapsed_time (ms) |
+|---|---|
+| 8 | 23.01 |
+| 16 | 25.72 |
+| 32 | 33.38 |
+| 64 | 65.98 |
+
+If you plot them using (hash_size)^2 and elapsed_time, you get almost a linear graph with a huge constant y-intercept:
+<img src="https://github.com/hj3yoo/mtg_card_detector/blob/master/figures/6_time_plot_1.png">
+
+Where is this fat constant of 22.4ms coming from? Well, we'd better look at how the [Imagehash library](https://github.com/JohannesBuchner/imagehash/blob/master/imagehash/__init__.py#L67-L72) deals with subtraction:
+```
+def __sub__(self, other):
+		if other is None:
+			raise TypeError('Other hash must not be None.')
+		if self.hash.size != other.hash.size:
+			raise TypeError('ImageHashes must be of the same shape.', self.hash.shape, other.hash.shape)
+		return numpy.count_nonzero(self.hash.flatten() != other.hash.flatten())
+```
+The code flattens both hashes before comparison. You might think, "How slow is that going to be?" Apparently fair amount.
+```
+a = np.ones([1, 1024], dtype=np.bool)
+start = time.time()
+for _ in range(10000):
+    if a is None:
+            raise TypeError('Other hash must not be None.')
+    if a.size != a.size:
+        raise TypeError('ImageHashes must be of the same shape.', self.hash.shape, other.hash.shape)
+    a.flatten()
+    a.flatten()
+end = time.time()
+elapsed = (end - start) * 1000
+print('%f' % elapsed)
+```
+
+The execution time of that code snippet on average is 11.65ms, which is slightly over half of 22.4ms of constant delay. That's a lot of time that can be cut out.
+By pre-emptively flattening the hashes and using the hash subtraction's code (yes I know it's not a good OOP design, but this is too much of a tradeoff), that constant time can be cut out significantly:
+
+| hash_size | elapsed_time (ms) |
+|---|---|
+| 8 | 9.9 |
+| 16 | 11.54 |
+| 32 | 18.55 |
+| 64 | 45.79 |
+
+Furthermore, turns out that hash size of 16 is sufficient enough to distinguish each cards in most of the case. Halving the hash size further knocked down 7-9ms, as you only need to compare about quarter of the bits compared to hash size of 32.
+
+------------------
+
+The other bottleneck is a something unfortunate. Turns out feeding the image through YOLO network consumes a constant 50 - 60ms per frame. Remember the processing time of (65+50)ms above? Yeah, that's where the 65ms is coming from.
+
+As hilarious and ironic it is, I would have to remove the network entirely to speed up the program... 
+**(((Facepalm into another dimension)))** 
+The program still works by replacing neural net with contour detection
+
+## Oct 13th, 2018
+
+Cleaning up everything to wrap up this project for now. If I can figure out how to move from bounding boxes of overlapping cards [(notes)](https://github.com/hj3yoo/mtg_card_detector#oct-4th-2018), I may come back to upgrade the project in the future. If you have any suggestion regarding this issue, please don't hesitate to let me know.
+
+Thank you for reading all the way up to here. Hope this project has helped you in some way.
\ No newline at end of file

--
Gitblit v1.10.0