We thought sigmoid activation function is best for more than 50 years. Just recently if you draw a straight diagonal line in the plot, we found that it significantly outperforms the other and it is called ReLU. Considering how long we have been doing wrong for half decades on a simple problem, how wrong is the current state of ML?