we've done our best to use the native provider's harness. all models were run on 'high' reasoning. this is still v1 and tons of room for improvement - really appreciate your feedback!