Support Vector Machines involve solving a large QP optimization problem... which is often done by running a variant of gradient descent (or one of the second- or quasi-second-order optimization algorithms, which involves finding the Hessian as well as the gradient)
Back when I was a researcher, I had a classification problem where the default random forest classifier built into scikitlearn completely dominated the neural method. The best numbers we got were something like a ~30% improvement from baseline for the neural network to ~95% for the random forest.
The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.
Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.
Support Vector Machines involve solving a large QP optimization problem... which is often done by running a variant of gradient descent (or one of the second- or quasi-second-order optimization algorithms, which involves finding the Hessian as well as the gradient)
I was under the impression that solving QP optimization problems (subject to constraints) was largely performed by SMO.
Back when I was a researcher, I had a classification problem where the default random forest classifier built into scikitlearn completely dominated the neural method. The best numbers we got were something like a ~30% improvement from baseline for the neural network to ~95% for the random forest.
The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.
Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.