Questions and Answers : Getting started & support : 1 (0x00000001) Unknown error code on M1
Message board moderation
| Author | Message |
|---|---|
rilianNew member Send message Joined: 28 Sep 26 Posts: 78 Credit: 5,428,405 RAC: 446,069 |
task failed after half day of calculation on M1 https://bitboinc.athena.org.tr/result.php?resultid=409309 <core_client_version>8.2.10</core_client_version> <![CDATA[ <message> process exited with code 1 (0x1, -255)</message> <stderr_txt> Keraunos Metal (bounded discrete-log search, Apple GPU) v1 bitboinc -- https://bitboinc.athena.org.tr/ coded by Alperen Yavuz kangaroo_run_metal: using Metal device "Apple M1" Keraunos Metal (bounded discrete-log search, Apple GPU) v1 bitboinc -- https://bitboinc.athena.org.tr/ coded by Alperen Yavuz kangaroo_run_metal: using Metal device "Apple M1" Keraunos Metal (bounded discrete-log search, Apple GPU) v1 bitboinc -- https://bitboinc.athena.org.tr/ coded by Alperen Yavuz kangaroo_run_metal: using Metal device "Apple M1" Keraunos Metal (bounded discrete-log search, Apple GPU) v1 bitboinc -- https://bitboinc.athena.org.tr/ coded by Alperen Yavuz kangaroo_run_metal: using Metal device "Apple M1" Keraunos Metal (bounded discrete-log search, Apple GPU) v1 bitboinc -- https://bitboinc.athena.org.tr/ coded by Alperen Yavuz kangaroo_run_metal: using Metal device "Apple M1" kangaroo_run_metal: command buffer error: Internal Error (0000000e:Internal Error) kangaroo_run_metal failed (device 0) -- no usable Metal device? 2026-09-30 09:20:11 (7682): called boinc_finish(1) </stderr_txt> ]]> I crunch for Ukraine
|
|
New member Send message Joined: 29 Sep 26 Posts: 9 Credit: 0 RAC: 0 |
It's great that you are contributing to the project, and we appreciate you sharing your results and effort! It is always encouraging to see volunteers dedicating significant time to the calculations. Based on the error logs you provided, the failure appears to be related to a Metal device issue on your M1, specifically noting a "command buffer error: Internal Error" and mentioning "no usable Metal device?". This type of problem is often a local environment configuration issue rather than a bug within the project itself. Since you are running the Metal build, I recommend checking the BOINC Manager's Messages tab to see if there are any specific warnings or messages about your GPU setup. Please also ensure that your system has the most current and correct Metal drivers and runtimes installed for your macOS environment, as a plan_class or driver mismatch is the most common cause for these kinds of GPU task failures. |
|
New member Send message Joined: 27 Sep 26 Posts: 83 Credit: 877,800 RAC: 67,962 |
I disabled Tilvar's response feature because it was giving very nonsensical answers; it will only handle moderation now.
|
|
New member Send message Joined: 29 Sep 26 Posts: 22 Credit: 0 RAC: 0 |
Found the real cause, and it was a genuine bug in the Metal (Apple GPU) code, not anything specific to your M1. The app creates a new Metal command buffer every round of the search, but the whole run only had one memory pool wrapping the entire multi-hour call -- so every round's GPU-side objects piled up unreleased instead of being cleaned up as it went. After enough rounds (which for you meant about half a day), Metal's internal resources ran out and the driver returned that generic Internal Error. Fixed by releasing each round's objects immediately instead of holding all of them for the whole run. Verified against a known-answer test on a real M1 before deploying. Rebuilt and redeployed for both Keraunos and Sphinx Metal (same bug, same fix, since Sphinx's Metal port was written from Keraunos's code). Sorry about the wasted half-day -- appreciate you posting the full stderr, that's what made this a quick diagnosis instead of a guess. |
rilianNew member Send message Joined: 28 Sep 26 Posts: 78 Credit: 5,428,405 RAC: 446,069 |
|
rilianNew member Send message Joined: 28 Sep 26 Posts: 78 Credit: 5,428,405 RAC: 446,069 |
|
rilianNew member Send message Joined: 28 Sep 26 Posts: 78 Credit: 5,428,405 RAC: 446,069 |
|
|
New member Send message Joined: 29 Sep 26 Posts: 22 Credit: 0 RAC: 0 |
@rilian: answering your three questions. New apps: already live. Keraunos Metal is now app_version 240 (v23) and Sphinx Metal is app_version 241 (v9), both with the fix -- no need to wait, just abort/re-download and BOINC will pick them up on the next scheduler contact. Does this affect Potamos too: no. I checked Potamos's Metal search code specifically -- it already creates and releases each round's GPU objects inside their own scope, so it never had the bug that hit Keraunos/Sphinx (a single memory pool wrapping the whole multi-hour run instead of one per round). Potamos Metal users don't need to do anything. Your M4 task: I checked, and it's currently running on the old app_version (157/v22), the one with the bug -- so yes, it's realistically at risk of the same Internal Error the longer it runs (your report showed it took about half a day on the M1). I'd abort it and let your host fetch a fresh task on the new version rather than risk losing 9h+ of work. On app size: noted, will look into trimming the Metal builds. |