It uses a descent with the following step:
Computationally heavy as you need to calculate the hessian all the time.
Why that step:
Setting this to zero gives:
Theorem:
Let
Remark:
This is exponentially faster than gradient descent. But needs more memory so it’s not used for high dimensions.
Proof:
NON EXAMINABLE!!!